Pith. sign in

REVIEW 4 major objections 4 minor 61 references

Vilya-2, a diffusion transformer that represents all molecules as atom-and-bond graphs, predicts the bound structures of chemically diverse peptides and small molecules with accuracy that far exceeds co-folding baselines, even when those ba

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:15 UTC pith:7F2JVWFH

load-bearing objection Vilya-2 is a real advance, but the CPSea training-date ambiguity could poison the headline Riptides number. the 4 major comments →

arxiv 2607.25156 v1 pith:7F2JVWFH submitted 2026-07-28 cs.LG

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

classification cs.LG
keywords all-atom representationdiffusion transformerprotein–peptide interface predictionconformer generationmacrocyclesconfidence calibrationsmall-molecule dockingfoundation model
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces Vilya-2, a diffusion transformer that represents peptides, small molecules, and protein targets identically as atom-and-bond graphs, with no residue tokens or molecule-type annotations. The authors claim this unified all-atom representation, combined with diverse pose sampling and a calibrated confidence score, lets the model predict bound peptide conformations far more accurately than co-folding models such as Boltz-2, even when those models are given the receptor crystal structure as a template. On the new Riptides benchmark, Vilya-2 recovers the peptide backbone within 2 Å RMSD in 59.1% of cases with 1000 samples, versus 40.9% for Boltz-2. The same architecture generalizes to small-molecule docking, macrocycles, and miniproteins larger than any training example, and can be fine-tuned to rank active compounds in drug-discovery campaigns.

Core claim

The paper's central claim is that a single all-atom, residue-free chemical representation—where a peptide, a macrocycle, a small molecule, and a protein are all atom-and-bond graphs—combined with ensemble sampling and a calibrated confidence score, is enough to accurately model protein–ligand interfaces across diverse chemistries. On the authors' new Riptides benchmark, Vilya-2 recovers the bound peptide backbone within 2 Å RMSD for 59.1% of cases when 1000 poses are sampled and ranked, compared with 40.9% for the co-folding baseline Boltz-2 even when Boltz-2 is given the receptor crystal structure as a template. The paper further claims that this advantage comes from both diversity in sampl

What carries the argument

The key machinery is the unified all-atom graph representation inherited from Vilya-1: the input is a chemical graph whose nodes are heavy atoms and whose edges are covalent bonds, with no residue tokens, molecule-type annotations, or multiple-sequence-alignment information. A diffusion transformer runs the diffusion process through the whole architecture and can generate diverse structural ensembles; a separately trained confidence head predicts per-atom lDDT from the chemical graph and predicted coordinates, and this plddt-ligand score ranks the sampled poses. For interface prediction, the model is additionally conditioned on a sparse distance matrix derived from receptor Cα and nucleic-ac

Load-bearing premise

The load-bearing premise is that the Riptides benchmark, manually curated by the authors from PDB entries after mid-2023 with unresolved residues and weak-contact entries removed, fairly represents the distribution of therapeutically relevant peptide interfaces and does not systematically favor Vilya-2 over co-folding baselines; if the curation biases the comparison, the reported accuracy gap could shrink or disappear.

What would settle it

Run Vilya-2 on all post-2023-06-01 PDB peptide–protein complexes that pass only automated filters (resolution < 3 Å, at least 20 atom contacts within 5 Å, receptor chain ≥ 50 residues) without the authors' manual exclusion of unresolved residues or weak-contact entries, comparing success rates to Boltz-2 under identical template conditioning; if the accuracy gap narrows to near parity, the benchmark curation rather than the model would be the source of the reported advantage.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Peptide therapeutics with non-canonical residues, macrocyclic topologies, and unusual covalent linkages can be structurally modeled without retraining or residue-level featurization, making them addressable by de novo design pipelines.
  • Inference-time scaling through diffusion sampling is a practical lever for interface accuracy: Vilya-2 keeps improving at 1000 samples on peptide targets, whereas co-folding baselines do not benefit from extra sampling.
  • A well-calibrated confidence score provides an absolute stop-criterion—designers can trust a pose with high plddt-ligand and know when more sampling is needed, rather than relying on relative ranking alone.
  • The same pretrained model transfers across small-molecule docking, conformer generation of large organics, and macrocycles, suggesting a single foundation model can serve multiple modalities in drug discovery.
  • Fine-tuning on fewer than a thousand experimental potency measurements yields roughly 3× enrichment for active compounds, indicating that the learned structural representations carry signal relevant to binding affinity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the all-atom representation delivers on its promise, the traditional separation between protein structure prediction, docking, and conformer generation may collapse into one modeling task, simplifying the software stack of structure-based drug design.
  • The receptor backbone conditioning means Vilya-2 assumes a known or reliably modeled target structure; a natural stress test is fully unbound receptor prediction with no structural prior, which is not covered by the reported benchmarks.
  • The PoseBusters prefilter—discarding physically invalid poses before confidence ranking—is a simple, general recipe that other generative structure models could adopt, and it suggests that physical-validity filters complement learned scoring rather than compete with it.
  • A testable extension would be to apply Vilya-2 to peptides with backbone N-methylation or ester linkages and to covalent inhibitors, to check whether the atom-graph representation handles bond-order and chirality variations beyond those in Riptides.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Vilya-2, a diffusion transformer that represents proteins, peptides, and small molecules uniformly as atomic graphs rather than residue-level tokens. It is trained for two tasks: single-molecule conformer generation and target-conditioned interface prediction, with a shared architecture later fine-tuned for confidence estimation, activity prediction, and property prediction. The central empirical claims are: (i) on a new, author-curated Riptides benchmark of 88 protein–peptide complexes, Vilya-2 recovers 54.1% of bound peptide backbones below 2 Å RMSD at 100 samples and 59.1% at 1000 samples, outperforming Boltz-2 even when Boltz-2 is given the receptor crystal structure as a template; (ii) Vilya-2 is state-of-the-art at small-molecule docking on PoseBusters, Runs N' Poses, and PoseX, including cross-docking; (iii) the conformer generator generalizes to macrocycles and disulfide-stapled miniproteins larger than training examples; and (iv) fine-tuned heads improve enrichment in an internal hit-to-lead campaign. The paper also releases the Riptides benchmark and benchmarking code.

Significance. If the results hold, Vilya-2 would be a meaningful advance: it demonstrates that a single all-atom representation can support peptide, small-molecule, macrocycle, and miniprotein modeling without residue-level or MSA-derived features, and that inference-time sampling with calibrated confidence scoring is effective for interface prediction. The paper includes useful controls: template-conditioned Boltz-2 baselines, Schrödinger docking comparisons, PoseBusters physical-validity filtering, explicit calibration plots, and sampling-scaling analysis. The release of the Riptides benchmark is a community asset. However, the headline generalization claims rest on a self-curated benchmark with an ambiguous training-data cutoff for the CPSea augmentation, and the reported confidence intervals are computed by resampling poses rather than benchmark targets. These issues are checkable and fixable, but they are load-bearing for the paper's central quantitative claims.

major comments (4)
  1. [§2.2] The training-data cutoff is ambiguous. The text says the dataset was expanded with protein–peptide complexes from CPSea and then states: 'training was restricted to PDB entries released on or before 2021–09–30.' It is not clear whether this restriction applies to the CPSea-derived complexes as well. This matters directly because Riptides — the benchmark used for the headline 59.1% result — is constructed from PDB entries released after 2023-06-01 and contains 51/88 non-canonical cyclic peptides, exactly the modality covered by CPSea. If CPSea complexes were not filtered by the 2021–09–30 cutoff, near-duplicates of training complexes can appear in the test set. Please (a) state explicitly whether the September 2021 cutoff was applied to CPSea-derived complexes; (b) if it was not, rerun the Riptides evaluation with any CPSea-derived training complexes released after 2021-09-30 excluded, or
  2. [§2.4] The reported '95% confidence intervals' are obtained by resampling 100 poses from a fixed pool of 1000 poses per target, then recomputing the success rate over the 88 targets. This procedure captures only pose-selection stochasticity; it does not reflect uncertainty about the benchmark success rate or about the difference between Vilya-2 and Boltz-2 across targets. For example, the Riptides 2 Å estimate of 54.1 ± 4.6% would have a substantially wider target-level binomial interval given n = 88. Please either relabel these intervals as within-system sampling variability, or add target-level bootstrap/binomial confidence intervals for the headline numbers and for the Vilya-2 vs. Boltz-2 comparison.
  3. [§3.1, Figure 1A] The text states: 'in all of these cases, none of the experimental structures we compare Vilya-2’s predictions against were present in the training data on which the model was trained.' At least one of the comparison structures, PDB 7S5G (the progenitor of Lipfendra), may have been released before the 2021–09–30 training cutoff; the manuscript does not provide release dates for the comparison structures. Please verify the release date of 7S5G and of the internal structures, and provide a table of release dates. If any comparison structure is in the training set, qualify the claim accordingly or exclude that comparison.
  4. [§2.3] The Riptides construction includes a manual curation step: entries with unresolved residues or 'inadequate contact' with the receptor are dropped. The criteria for 'inadequate contact' are not quantified, and the manual step is performed by the authors who are also proposing the benchmark and evaluating their own model. Because this benchmark is the primary evidence for the central claim of peptide-modeling superiority, please (a) release the full list of candidate entries with the reasons for exclusion, (b) define the contact criterion in terms of a reproducible threshold (e.g., number of atom pairs, buried surface area), and (c) report whether the Vilya-2 vs. Boltz-2 comparison is stable under alternative inclusion thresholds or when the manually excluded entries are included.
minor comments (4)
  1. [Abstract] There is a typo in the final sentence: 'the design and evaluation of de novo of peptide therapeutics' should read 'de novo peptide therapeutics'.
  2. [§2.2] The text uses 'released on or before 2021–09–30' for the PDB restriction but does not define whether 'released' means PDB release date or deposition date. Please state which date is used and apply the same convention consistently.
  3. [Table 1] The columns 'PB-Valid = True' and 'PB-Valid prefilter' are clear from the text but would benefit from a one-line definition in the table caption: the former excludes poses that fail PoseBusters checks, while the latter discards invalid poses before confidence ranking so that the output always passes all checks.
  4. [§2.3] The Riptides selection criteria refer to 'at least 20 unique atom pairs with a distance < 5 Å between the ligand and its closest protein receptor chain.' It would be useful to state whether this includes only heavy atoms and how symmetry-related copies of the ligand were handled.

Circularity Check

0 steps flagged

No significant circularity: Vilya-2's structural predictions are evaluated on held-out external benchmarks, and the Vilya-1 self-citation supplies architecture and protocol rather than substituting for evidence.

full rationale

The main derivation chain — train a graph-based diffusion transformer on PDB/CPSea structures, generate diverse poses, rank with a separately trained lDDT confidence head, and measure RMSD on externally released benchmarks — does not reduce any evaluated quantity to a fitted input. Riptides entries are selected from PDB releases after 2023-06-01, while the paper states that training was restricted to PDB entries released on or before 2021-09-30; the confidence model's calibration is evaluated on held-out benchmark distributions and is not used to fit the structural model. The Vilya-1 self-citation supplies the architecture, benchmarks, and fine-tuning protocol, but Vilya-2's central interface-prediction claims are supported by external comparisons (PoseBusters, Runs N' Poses, PoseX, and Boltz-2/Schrödinger baselines) rather than by appeal to Vilya-1's own conclusions. One missing-support caveat is worth flagging: Section 2.2's temporal cutoff is stated only for 'PDB entries released on or before 2021–09–30,' and the preceding CPSea augmentation sentence does not explicitly confirm that CPSea-derived complexes were also filtered to pre-2021-09-30. Since Riptides contains 51/88 non-canonical cyclic peptides, a post-cutoff CPSea corpus could in principle overlap the benchmark, which would undermine the headline generalization claim. This is a data-leakage/correctness concern, not a demonstrated by-construction equivalence, so it does not raise the circularity score above 1.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

This is an empirical machine learning paper; no new physical entities are introduced. The central claim rests on standard domain assumptions about PDB ground truth, the validity of the all-atom representation, and the fairness of the self-built benchmark.

free parameters (2)
  • Interaction crop size = 1024 atoms
    Chosen to cover ~120 residue domain; limits model to local interaction region, potentially excluding long-range contacts.
  • Crop jitter median radial displacement = 3.2 Å
    Chosen to prevent trivial memorization of ligand location; affects diversity of training crops.
axioms (5)
  • domain assumption PDB crystal structures are accurate ground truth for molecular interfaces.
    Used in training and evaluation; errors in deposited structures could bias both training and benchmark assessment.
  • domain assumption The diffusion training objective plus auxiliary losses (distogram, stereochemical) yields a well-calibrated generative model over molecular conformations.
    This underpins the model's ability to sample diverse poses; not formally proven.
  • domain assumption The local, receptor-conditioned formulation (known backbone coordinates) is a valid and practical approximation for drug discovery scenarios.
    The paper argues most campaigns know the target structure, but this may not hold for all novel targets.
  • ad hoc to paper The manual curation of the Riptides benchmark does not introduce systematic bias favoring Vilya-2.
    Authors manually dropped entries with unresolved residues or inadequate contact; without an external audit, this could bias results.
  • domain assumption The plddt-ligand confidence score is transferable to novel chemistries.
    Confidence model is trained on the model's own outputs; calibration is shown on benchmarks, but generalization to diverse non-canonical peptides is assumed.

pith-pipeline@v1.3.0-alltime-deepseek · 16208 in / 14741 out tokens · 131103 ms · 2026-08-01T03:15:47.762729+00:00 · methodology

0 comments
read the original abstract

Structure-prediction networks built on co-evolutionary statistics have transformed protein-based drug discovery, yet their accuracy does not extend to peptide therapeutics--an increasingly important modality defined by non-canonical residues, macrocyclization, and complex topologies. We introduce Vilya-2, a diffusion transformer that extends the all-atom representation of Vilya-1 from modeling individual molecules to modeling their interactions with protein targets. This all-atom representation enables transfer learning between different molecular types, and delivers highly accurate structural modeling of peptides across sizes, classes, and compositions bound to therapeutically relevant targets. By generating diverse structural ensembles and ranking them with calibrated confidence, Vilya-2 recovers 59.1% of peptide interfaces to sub-2 {\AA} backbone RMSD, far exceeding the performance of a representative co-folding model even when that model is given the bound receptor as a template. In addition, Vilya-2 is state-of-the-art at small-molecule docking, and generalizes to novel protein-small molecule complexes unlike those seen in training. It also generalizes to modeling molecular conformations of diverse macrocycles and disulfide-stapled miniproteins several-fold larger than any molecule seen in training. Finally, Vilya-2 can be used as a foundation model, and fine-tuned to enrich for active compounds in hit-to-lead campaigns. By unifying predictive accuracy with broad generalizability across chemical space, Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Figures

Figures reproduced from arXiv: 2607.25156 by Adam P. Moyer, Benjamin D. Sellers, Chase A. P. Wood, CJ San Felipe, Ivan Anishchanka, Jeffrey K. Holden, Milad Salem, Naozumi Hiranuma, Patrick J. Salveson, Stephen Rettie, Vilya Research: Pascal Sturmfels.

Figure 1
Figure 1. Figure 1: Accurate prediction of protein-ligand interactions with Vilya-2. A) Structural superpositions of Vilya-2 interaction predictions (green) with reference X-ray crystal structures (gray). Examples include two macrocyclic drug candidates (daraxonrasib and zolucatetide) and two approved macrocyclic drug (Lipfendra and Icotyde). Predictions for daraxonrasib and Icotyde are superimposed on their respective co-cry… view at source ↗
Figure 2
Figure 2. Figure 2: An overview of Vilya-2. Structure-prediction networks form the core of Vilya-2: a conformer generator and a target-conditioned complex predictor, which generate ensembles of conformers or of protein–ligand complexes respectively. The two share the same architecture, input features, and training losses, and are trained first. Their architecture is then reused for specialized heads, which are fine-tuned to p… view at source ↗
Figure 3
Figure 3. Figure 3: Composition of the Riptides benchmark. A) Breakdown of the 88 entries in the benchmark by topology (cyclic vs linear) and chemical composition (canonical vs non-canonical) with cell shading proportional to count; non-canonical cyclic peptides dominate the set (51 of 88). B) Distribution of ligand size, in number of atoms. C) Distribution of ligand length, in number of residues; macrocycles annotated as non… view at source ↗
Figure 4
Figure 4. Figure 4: Interplay between efficient sampling and accurate scoring drives Vilya-2 predictive accuracy. A) Increased sampling combined with pLDDT-based ranking yields higher prediction success. The fraction of targets with a successfully predicted bound ligand conformation (RMSD < 2˚A) is plotted as a function of the number of generated samples. Success rates are evaluated using top-ranked selection based on plddt-l… view at source ↗
Figure 5
Figure 5. Figure 5: Activity Prediction with Vilya-2 on an internal hit-to-lead optimization campaign. A) Fold enrichment in hit rate among the top 20% of ranked compounds (EF20%) for two ligand-only baselines—K-nearest neighbors (KNN) on ECFP fingerprints and ChemProp—and the two Vilya-2 variants. The conformer-based model (light green hatched), fine-tuned on ligand conformers alone, and the complex-based model (light green)… view at source ↗
Figure 6
Figure 6. Figure 6: Vilya-2 accurately recapitulates structures of disulfide-stapled miniproteins. A,B) Overlays of experimental NMR structures (gray cartoon) with the best-sampled Vilya-2 conformation (green sticks; n = 4,000 generated samples). For clarity, Vilya-2 models are displayed using ring scaffold heavy atoms only, with side chains omitted. Asterisks (*) denote structures containing non-canonical amino acids, includ… view at source ↗
Figure 7
Figure 7. Figure 7: Vilya-2 accurately recapitulates structures of large organic molecules. Overlays of X-ray crystal structures from the Cambridge Structural Database (gray) with Vilya-2 models (green). All depicted molecules contain > 128 heavy atoms. For each molecule, the model with the lowest all-atom RMSD to the crystal structure from an ensemble of 1,200 generated samples is shown. All displayed Vilya-2 models achieve … view at source ↗
Figure 8
Figure 8. Figure 8: Conformer sampling success rates on a rigorous benchmark of X-ray characterized macrocycles. Both Vilya-2 variants exhibit superior structural prediction performance compared to Vilya-1, with particularly pronounced gains in the high-accuracy regime (dark bars; Ring RMSD < 0.5˚A) compared to standard accuracy thresholds (light bars). Vilya Research 13 [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Vilya-2 property prediction results. Fold enrichment of favorable compounds in the top 20% of predictions (EF20%) for the Vilya-2 conformer generator fine-tuned as a property-prediction head (green), compared with two ligand-only baselines—K-nearest neighbors (KNN) on ECFP fingerprints (hatched) and ChemProp (gray)—on chromatographic logD (ChromLogD), MDCK permeability, PAMPA permeability, and kinetic solu… view at source ↗
Figure 10
Figure 10. Figure 10: Comprehensive architectural optimization yields a highly efficient and accurate conformer generator. A) Empirical Pareto frontier mapping high-resolution structural accuracy against computational cost. The fraction of macrocycles achieving < 0.5˚A Ring RMSD is plotted as a function of inference FLOPs per sample. The Pareto front (green line) is derived from a systematic architectural sweep of 112 trained … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 2 canonical work pages

  1. [1]

    MJ Gooderham et al. “LB1142 Phase 3 results from an innovative trial design of treating plaque psoriasis involving difficult-to-treat, high-impact sites with icotrokinra, a targeted oral peptide that selectively inhibits the IL-23–receptor”. In:Journal of Investigative Dermatology145.8 (2025), S199. Vilya Research 15

  2. [2]

    A placebo-controlled trial of the oral PCSK9 inhibitor enlicitide

    Ann Marie Navar et al. “A placebo-controlled trial of the oral PCSK9 inhibitor enlicitide”. In:New England Journal of Medicine394.6 (2026), pp. 529–539

  3. [3]

    Klempner et al

    Samuel J. Klempner et al. “A phase 1/2 trial of FOG-001, a first-in-class direct β-catenin:TCF4 inhibitor, preliminary safety and efficacy in patients with solid tumors bearing Wnt pathway-activating mutations (WPAM+)”. In:Proceedings of the AACR-NCI-EORTC International Conference on Molecular Targets and Cancer Therapeutics. Vol. 24. 10 Suppl. Abstract n...

  4. [4]

    Chemical optimization of an orally bioavailable macrocyclic peptide KRAS inhibitors that achieve KRAS selectivity by recognizing a single amino acid difference

    Mirai Kage et al. “Chemical optimization of an orally bioavailable macrocyclic peptide KRAS inhibitors that achieve KRAS selectivity by recognizing a single amino acid difference”. In:Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts). Vol. 86. 7 Suppl. Abstract nr 5138. San Diego, CA: AACR, Apr. 2026

  5. [5]

    Daraxonrasib plus chemotherapy (CT) as first-line (1L) treatment for patients (Pts) with metastatic pancreatic adenocarcinoma (mPDAC)

    Brian M. Wolpin et al. “Daraxonrasib plus chemotherapy (CT) as first-line (1L) treatment for patients (Pts) with metastatic pancreatic adenocarcinoma (mPDAC)”. In:Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts). Vol. 86. 8 Suppl. Abstract nr LB407. San Diego, CA...

  6. [6]

    Orally Bioavailable Cyclin A/B RxL Inhibitors: Optimization of a Novel Class of Macrocyclic Peptides That Target E2F-High and G1–S-Checkpoint-Compromised Cancers

    Justin A Shapiro et al. “Orally Bioavailable Cyclin A/B RxL Inhibitors: Optimization of a Novel Class of Macrocyclic Peptides That Target E2F-High and G1–S-Checkpoint-Compromised Cancers”. In:Journal of Medicinal Chemistry 69.5 (2026), pp. 5441–5460

  7. [7]

    Early clinical activity from the phase 1 evaluation of CID-078, a novel Cyclin A/B-RxL inhibitor, in patients with advanced solid tumors

    Afshin Dowlati et al. “Early clinical activity from the phase 1 evaluation of CID-078, a novel Cyclin A/B-RxL inhibitor, in patients with advanced solid tumors”. In:Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts). Vol. 86. 8 Suppl. Abstract nr CT023. San Diego, ...

  8. [8]

    Therapeutic peptides: current applications and future directions

    Lei Wang et al. “Therapeutic peptides: current applications and future directions”. In:Signal transduction and targeted therapy7.1 (2022), p. 48

  9. [9]

    Pearl: A Foundation Model for Placing Every Atom in the Right Location

    Alejandro Dobles et al. “Pearl: A Foundation Model for Placing Every Atom in the Right Location”. In:arXiv preprint arXiv:2510.24670(2025)

  10. [10]

    Boltz-1 democratizing biomolecular interaction modeling

    Jeremy Wohlwend et al. “Boltz-1 democratizing biomolecular interaction modeling”. In:BioRxiv(2025), pp. 2024– 11

  11. [11]

    Chai-1: Decoding the molecular interactions of life

    Chai Discovery team et al. “Chai-1: Decoding the molecular interactions of life”. In:BioRxiv(2024), pp. 2024–10

  12. [12]

    The past, present and future of de novo protein design

    Wei Yang et al. “The past, present and future of de novo protein design”. In:Nature652.8112 (2026), pp. 1139– 1152

  13. [13]

    De novo design of all-atom biomolecular interactions with rfdiffusion3

    Jasper Butcher et al. “De novo design of all-atom biomolecular interactions with rfdiffusion3”. In:bioRxiv(2025)

  14. [14]

    Boltzdesign1: Inverting all-atom structure prediction model for generalized biomolecular binder design

    Yehlin Cho et al. “Boltzdesign1: Inverting all-atom structure prediction model for generalized biomolecular binder design”. In:BioRxiv(2025), pp. 2025–04

  15. [15]

    Zero-shot antibody design in a 24-well plate

    Chai Discovery Team et al. “Zero-shot antibody design in a 24-well plate”. In:bioRxiv(2025), pp. 2025–07

  16. [16]

    De novo design of high-affinity protein binders with AlphaProteo

    Vinicius Zambaldi et al. “De novo design of high-affinity protein binders with AlphaProteo”. In:arXiv preprint arXiv:2409.08022(2024)

  17. [17]

    Accurate structure prediction of biomolecular interactions with AlphaFold 3

    Josh Abramson et al. “Accurate structure prediction of biomolecular interactions with AlphaFold 3”. In:Nature 630.8016 (2024), pp. 493–500

  18. [18]

    Improving de novo protein binder design with deep learning

    Nathaniel R Bennett et al. “Improving de novo protein binder design with deep learning”. In:Nature Communications 14.1 (2023), p. 2625

  19. [19]

    Evaluating generalization in protein–ligand cofolding methods

    Peter ˇSkrinjar et al. “Evaluating generalization in protein–ligand cofolding methods”. In:Nature Structural & Molecular Biology(2026), pp. 1–13

  20. [20]

    Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions

    Matthew R Masters, Amr H Mahmoud, and Markus A Lill. “Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions”. In:Nature Communications16.1 (2025), p. 8854

  21. [21]

    AlphaFold3 for noncanonical cyclic peptide modeling: hierarchical benchmarking reveals accuracy and practical guidelines

    Chengyun Zhang et al. “AlphaFold3 for noncanonical cyclic peptide modeling: hierarchical benchmarking reveals accuracy and practical guidelines”. In:Journal of Chemical Information and Modeling65.18 (2025), pp. 9777–9789

  22. [22]

    Beyond 20 in the 21st century: prospects and challenges of non-canonical amino acids in peptide drug discovery

    Jennifer L Hickey et al. “Beyond 20 in the 21st century: prospects and challenges of non-canonical amino acids in peptide drug discovery”. In:ACS Medicinal Chemistry Letters14.5 (2023), pp. 557–565

  23. [23]

    Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

    Vilya Research et al. “Vilya-1: An all-atom foundation model for macrocycle structure prediction and design”. In: arXiv preprint arXiv:2607.09998(2026). arXiv:2607.09998 [cs.LG]. Vilya Research 16

  24. [24]

    Generalized biomolecular modeling and design with RoseTTAFold All-Atom

    Rohith Krishna et al. “Generalized biomolecular modeling and design with RoseTTAFold All-Atom”. In:Science 384.6693 (2024), eadl2528

  25. [25]

    Modeling protein–small molecule conformational ensembles with PLACER

    Ivan Anishchenko et al. “Modeling protein–small molecule conformational ensembles with PLACER”. In:Proceedings of the National Academy of Sciences122.45 (2025), e2427161122

  26. [26]

    Triangle Multiplication Is All You Need For Biomolecular Structure Representations

    Jeffrey Ouyang-Zhang et al. “Triangle Multiplication Is All You Need For Biomolecular Structure Representations”. In:arXiv preprint arXiv:2510.18870(2025)

  27. [27]

    Proteina: Scaling flow-based protein structure generative models

    Tomas Geffner et al. “Proteina: Scaling flow-based protein structure generative models”. In:arXiv preprint arXiv:2503.00710(2025)

  28. [28]

    CPSea: large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design

    Ziyi Yang et al. “CPSea: large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design”. In:Advances in Neural Information Processing Systems38 (2026)

  29. [29]

    Inherent versus induced protein flexibility: Comparisons within and between apo and holo structures

    Jordan J Clark et al. “Inherent versus induced protein flexibility: Comparisons within and between apo and holo structures”. In:PLoS computational biology15.1 (2019), e1006705

  30. [30]

    ECOD: an evolutionary classification of protein domains

    Hua Cheng et al. “ECOD: an evolutionary classification of protein domains”. In:PLoS computational biology10.12 (2014), e1003926

  31. [31]

    lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests

    Valerio Mariani et al. “lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests”. In:Bioinformatics29.21 (2013), pp. 2722–2728

  32. [32]

    Lora: Low-rank adaptation of large language models

    Edward J Hu et al. “Lora: Low-rank adaptation of large language models.” In:Iclr1.2 (2022), p. 3

  33. [33]

    Attention-based deep multiple instance learning

    Maximilian Ilse, Jakub Tomczak, and Max Welling. “Attention-based deep multiple instance learning”. In:Interna- tional conference on machine learning. PMLR. 2018, pp. 2127–2136

  34. [34]

    HTRF: a technology tailored for drug discovery–a review of theoretical aspects and recent applications

    Franc ¸ois Degorce et al. “HTRF: a technology tailored for drug discovery–a review of theoretical aspects and recent applications”. In:Current chemical genomics3 (2009), p. 22

  35. [35]

    SIMPD: an algorithm for generating simulated time splits for validating machine learning approaches

    Gregory A Landrum et al. “SIMPD: an algorithm for generating simulated time splits for validating machine learning approaches”. In:Journal of cheminformatics15.1 (2023), p. 119

  36. [36]

    PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences

    Martin Buttenschoen, Garrett M Morris, and Charlotte M Deane. “PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences”. In:Chemical Science15.9 (2024), pp. 3130–3139

  37. [37]

    Posex: Ai defeats physics approaches on protein-ligand cross docking

    Yize Jiang et al. “Posex: Ai defeats physics approaches on protein-ligand cross docking”. In:arXiv preprint arXiv:2505.01700(2025)

  38. [38]

    Avoidable errors in deposited macromolecular structures: an impediment to efficient data mining

    Zbigniew Dauter et al. “Avoidable errors in deposited macromolecular structures: an impediment to efficient data mining”. In:IUCrJ1.3 (2014), pp. 179–193

  39. [39]

    Models of protein–ligand crystal structures: trust, but verify

    Marc C Deller and Bernhard Rupp. “Models of protein–ligand crystal structures: trust, but verify”. In:Journal of computer-aided molecular design29.9 (2015), pp. 817–836

  40. [40]

    Detect, correct, retract: How to manage incorrect structural models

    Alexander Wlodawer et al. “Detect, correct, retract: How to manage incorrect structural models”. In:The FEBS journal285.3 (2018), pp. 444–466

  41. [41]

    Boltz-2: Towards accurate and efficient binding affinity prediction

    Saro Passaro et al. “Boltz-2: Towards accurate and efficient binding affinity prediction”. In:BioRxiv(2025)

  42. [42]

    Protein and ligand preparation: parameters, protocols, and influence on virtual screening enrichments

    G Madhavi Sastry et al. “Protein and ligand preparation: parameters, protocols, and influence on virtual screening enrichments”. In:Journal of computer-aided molecular design27.3 (2013), pp. 221–234

  43. [43]

    Schr ¨odinger, LLC, New York, NY

    Schr¨odinger, LLC.Schr ¨odinger Release 2026-1: Protein Preparation Workflow; Epik; Impact; Prime. Schr ¨odinger, LLC, New York, NY. Protein Preparation Workflow, Schr ¨odinger Release 2026-1; Epik (2024); Impact; Prime (2025). New York, NY, 2026

  44. [44]

    NeuralPLexer3: accurate biomolecular complex structure prediction with flow models

    Jarren Zhuoran Qiao et al. “NeuralPLexer3: accurate biomolecular complex structure prediction with flow models”. In:Advances in Neural Information Processing Systems38 (2026), pp. 1–38

  45. [45]

    Accurate physics-based flexible docking of macrocyclic ligands

    Jacob Robson-Tull and Jo ˜ao PGLM Rodrigues. “Accurate physics-based flexible docking of macrocyclic ligands”. In: Journal of Medicinal Chemistry69.4 (2026), pp. 4745–4754

  46. [46]

    Glide: a new approach for rapid, accurate docking and scoring. 1. Method and assessment of docking accuracy

    Richard A Friesner et al. “Glide: a new approach for rapid, accurate docking and scoring. 1. Method and assessment of docking accuracy”. In:Journal of medicinal chemistry47.7 (2004), pp. 1739–1749

  47. [47]

    De novo mapping of α-helix recognition sites on protein surfaces using unbiased libraries

    Kunhua Li et al. “De novo mapping of α-helix recognition sites on protein surfaces using unbiased libraries”. In: Proceedings of the National Academy of Sciences119.52 (2022), e2210435119. Vilya Research 17

  48. [48]

    JNJ-77242113, a highly potent, selective peptide targeting the IL-23 receptor, provides robust IL-23 pathway inhibition upon oral dosing in rats and humans

    Anne M. Fourie et al. “JNJ-77242113, a highly potent, selective peptide targeting the IL-23 receptor, provides robust IL-23 pathway inhibition upon oral dosing in rats and humans”. In:Scientific Reports14.1 (July 2024), p. 17515.doi:10.1038/s41598-024-67371-5.url:https://doi.org/10.1038/s41598-024-67371-5

  49. [49]

    Factors governing helical preference of peptides containing multiple alpha, alpha-dialkyl amino acids

    Garland R Marshall et al. “Factors governing helical preference of peptides containing multiple alpha, alpha-dialkyl amino acids.” In:Proceedings of the National Academy of Sciences87.1 (1990), pp. 487–491

  50. [50]

    James Cregg et al. “Discovery of daraxonrasib (RMC-6236), a potent and orally bioavailable RAS (ON) multi- selective, noncovalent tri-complex inhibitor for the treatment of patients with multiple RAS-addicted cancers”. In: Journal of Medicinal Chemistry68.6 (2025), pp. 6064–6083

  51. [51]

    Series of novel and highly potent cyclic peptide PCSK9 inhibitors derived from an mRNA display screen and optimized via structure-based design

    Candice Alleyne et al. “Series of novel and highly potent cyclic peptide PCSK9 inhibitors derived from an mRNA display screen and optimized via structure-based design”. In:Journal of medicinal chemistry63.22 (2020), pp. 13796– 13824

  52. [52]

    A series of novel, highly potent, and orally bioavailable next-generation tricyclic peptide PCSK9 inhibitors

    Thomas J Tucker et al. “A series of novel, highly potent, and orally bioavailable next-generation tricyclic peptide PCSK9 inhibitors”. In:Journal of medicinal chemistry64.22 (2021), pp. 16770–16800

  53. [53]

    Inference-time scaling for complex tasks: Where we stand and what lies ahead

    Vidhisha Balachandran et al. “Inference-time scaling for complex tasks: Where we stand and what lies ahead”. In: arXiv preprint arXiv:2504.00294(2025)

  54. [54]

    Isomorphic Labs Team.Accurate Predictions of Novel Biomolecular Interactions with IsoDDE. Feb. 2026.doi: 10.5281/zenodo.19699685

  55. [55]

    The coming of age of de novo protein design

    Po-Ssu Huang, Scott E Boyken, and David Baker. “The coming of age of de novo protein design”. In:Nature 537.7620 (2016), pp. 320–327

  56. [56]

    Chemprop: a machine learning package for chemical property prediction

    Esther Heid et al. “Chemprop: a machine learning package for chemical property prediction”. In:Journal of Chemical Information and Modeling64.1 (2023), pp. 9–17

  57. [57]

    Accurate de novo design of hyperstable constrained peptides

    Gaurav Bhardwaj et al. “Accurate de novo design of hyperstable constrained peptides”. In:Nature538.7625 (2016), pp. 329–335

  58. [58]

    Heterogeneous-Backbone Foldamer Mimics of a Computationally Designed, Disulfide-Rich Miniprotein

    Chino C Cabalteja, Daniel S Mihalko, and W Seth Horne. “Heterogeneous-Backbone Foldamer Mimics of a Computationally Designed, Disulfide-Rich Miniprotein”. In:ChemBioChem20.1 (2019), pp. 103–110

  59. [59]

    Heterogeneous-backbone proteomimetic analogues of lasiocepsin, a disulfide-rich antimi- crobial peptide with a compact tertiary fold

    Chino C Cabalteja et al. “Heterogeneous-backbone proteomimetic analogues of lasiocepsin, a disulfide-rich antimi- crobial peptide with a compact tertiary fold”. In:ACS chemical biology17.4 (2022), pp. 987–997

  60. [60]

    MDCK (Madin–Darby canine kidney) cells: a tool for membrane permeability screening

    Jennifer D Irvine et al. “MDCK (Madin–Darby canine kidney) cells: a tool for membrane permeability screening”. In:Journal of pharmaceutical sciences88.1 (1999), pp. 28–33

  61. [61]

    https://github.com/NVIDIA/cuEquivariance

    NVIDIA.cuEquivariance. https://github.com/NVIDIA/cuEquivariance. Version 0.10.0, accessed 2026-07-09. 2025. A Supplement Table S1:Ring RMSD of Vilya-2 predictions for the 21 disulfide-stapled miniproteins from three studies [57, 58, 59]. For each miniprotein (PDB ID), the value is the lowest ring RMSD ( ˚A) between any of the n = 4,000 generated Vilya-2 c...