Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

Vilya-1 samples near-native macrocycle ring shapes across arbitrary chemistries at roughly double the success rate of physics-based methods, and transfers that skill to property prediction and design.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 01:08 UTC pith:54C5UA32

load-bearing objection Real sampling gains on macrocycles with a clean all-atom design, but the headline numbers rest on proprietary training mixtures that outsiders cannot audit. the 4 major comments →

arxiv 2607.09998 v1 pith:54C5UA32 submitted 2026-07-10 cs.LG q-bio.BM

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

classification cs.LG q-bio.BM
keywords macrocyclic peptidesconformer generationall-atom representationmembrane permeabilityfoundation modelnon-canonical amino acidsdiffusion modelsdrug design
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Macrocyclic peptides sit between small molecules and proteins: they can bind hard protein surfaces yet still aim for cell entry and oral dosing. Existing structure tools either assume protein-like residues, struggle with ring flexibility, or work only for narrow chemistries, so design campaigns still rely on chemistry-specific sampling scripts. This paper introduces Vilya-1, a single all-atom deep learning model that takes only chemical atom and bond features, runs a unified diffusion process, and is trained on mixed crystal and computed structures spanning peptides, non-peptidic macrocycles, and small molecules. On held-out X-ray, NMR, and receptor-bound macrocycles it recovers near-native ring geometries far more often than physics-based samplers, co-folding networks, and prior deep-learning conformer generators, with especially large gains when non-canonical residues appear. The same backbone is fine-tuned for confidence ranking, multi-property prediction including membrane permeability, and discrete redesign that preserves binding motifs while changing size, chemistry, and topology. A sympathetic reader cares because accurate, chemistry-agnostic conformational sampling is the bottleneck that has limited computational design of drug-like macrocycles.

Core claim

Vilya-1, an all-atom diffusion model with a uniform chemical feature set, samples near-native macrocycle ring conformations (ring RMSD under 1 Å) in 89.2% of cases on a 66-structure X-ray cyclic-peptide test set—more than double the success rate of leading physics- and knowledge-based methods and several times that of co-folding networks and prior deep-learning conformer generators—while remaining accurate on NMR ensembles, receptor-bound poses, non-canonical chemistries, and small molecules, and transferring via fine-tuning to confidence estimation, developability property prediction, and motif-preserving generative design.

What carries the argument

A unified all-atom equivariant transformer with a single diffusion path (no separate structure trunk), atom-level scalar, pair, and vector features that omit residue type and positional encodings, trained with diffusion, distogram, and tetrahedral-chirality losses; this representation lets one network cover peptides, mixed, and non-peptidic systems and is fine-tuned for confidence and multi-property heads.

Load-bearing premise

That training on public crystal structures plus proprietary computed macrocycle ensembles teaches real structural principles rather than patterns that only hold because test molecules still resemble that training mix.

What would settle it

Assemble a new panel of macrocycles whose ring scaffolds, cyclization chemistries, and non-canonical building blocks have near-zero fingerprint and topological overlap with both the public crystal training data and the proprietary computational ensembles; if ring-RMSD success under 1 Å collapses to the level of physics-based samplers, the generalization claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Drug-discovery teams can score and rank macrocycle analogs for permeability, hydrophobicity, and solubility without writing chemistry-specific conformational protocols.
  • Display-derived peptide hits can be computationally miniaturized into smaller rings while holding the binding motif in place.
  • Ligand-only energy landscapes from Vilya-1 ensembles plus machine-learned potentials can enrich for binders without modeling the protein target.
  • Design can freely use thioether, γ-lactam, stapled-helix, and other non-head-to-tail topologies that match common high-throughput display formats.
  • The same sampler can serve as a general small-molecule conformer generator when macrocycle specificity is not required.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If residue-level tokenization is the main reason co-folding networks degrade on non-canonicals, the same atom-level diffusion pretraining may improve protein–ligand co-folding for exotic ligands more broadly.
  • The largest property gains appear inside closely related analog series, suggesting the practical deployment mode is series-specific fine-tuning rather than one global property model.
  • Near-independence of accuracy from Tanimoto similarity to training molecules implies useful transfer to other constrained flexible systems such as medium natural products and flexible linkers.
  • Because the model already samples near local energy minima, coupling it to cheaper confidence ranking may displace routine DFT-level rescoring for early macrocycle triage.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The manuscript introduces Vilya-1, an all-atom diffusion model for macrocycle conformer generation that uses a unified chemical feature set (no residue-type or positional encodings) and is trained on mixed public crystal structures plus computationally generated macrocycle ensembles. On a 66-structure X-ray cyclic-peptide test set it reports 89.2% success at ring RMSD < 1 Å (100 samples), roughly doubling Schrödinger Prime-MCS and RDKit ETKDGv3 and far exceeding Boltz-2, RF3, and several DL conformer generators; similar advantages are claimed on NMR ensembles, receptor-bound macrocycles, and small-molecule bound poses. Fine-tuned confidence and multi-property heads, plus discrete design demos (motif re-looping, non-head-to-tail topologies, Pnear enrichment), are presented as evidence that Vilya-1 functions as a foundation model for macrocycle drug discovery.

Significance. If the sampling gains hold under stricter leakage controls, this would be a genuine advance for a therapeutically important chemical class where physics-based sampling is slow and co-folding networks degrade on non-canonical chemistry. Strengths include multiple independent experimental test sets with an explicit, stringent ring-scaffold RMSD definition; a clear architectural bet on atom-level tokenization; PoseBusters/Tanimoto checks on held-out CSD small molecules; time-based internal property splits; and concrete design use-cases (thioether/γ-lactam/stapled topologies) that existing protein-centric tools struggle with. The work is industrially relevant and, if reproducible, would shift practice for macrocycle conformational prioritization and hit miniaturization.

major comments (4)
  1. Methods §2.2–2.3 and Results §3.1: The headline 89.2% X-ray success (and the NMR/bound gains) rest on generalization, yet training mixes public crystals with proprietary computationally generated macrocycle ensembles whose generators, accuracy, and chemical coverage are not described. Exclusion is only >70% Morgan-Tanimoto to test molecules. That filter does not rule out near-identical ring scaffolds, stereochemical variants, or imprint from the same physics generators later used as baselines. The only memorization analysis (Fig. 3A) is on CSD small molecules, not on the 66 X-ray / 260 NMR / 240 bound macrocycle sets that drive the central claim. Please (i) quantify training-set composition by source and size class, (ii) report max scaffold/ring-fingerprint similarity of each test macrocycle to any training structure (including computational ones), and (iii) show success vs. that similar
  2. Methods §2.2 (Training) and §2.4 (Evaluation): For the 66 X-ray and 260 NMR sets, state explicitly whether any PDB or CSD entry used as a test structure (or a close computational analogue of it) entered pretraining via the internal ensembles. The paper already asserts that Fig. 1A examples and the 240 bound ligands were unseen; the same statement and supporting similarity tables are needed for the X-ray and NMR test sets that produce the largest reported gaps. If computational training structures were produced by protocols related to Prime-MCS or ETKDG, also discuss whether that creates an unfair comparison when those methods are the baselines.
  3. Results §3.2 / Fig. 5: Property-prediction gains are largest on proprietary internal series (time-split), where structure-aware pretraining helps resolve close analogs. External PAMPA enrichment is closer to ChemProp/CheMeleon. Because internal labels and series membership are not shareable, the foundation-model claim for developability needs either (a) a fully external, scaffold-split multi-property benchmark with released splits, or (b) a stronger ablation showing that the conformer-conditioned head (not just the pretrained trunk) is required for the internal EF gains. As written, the internal results are suggestive but not independently verifiable.
  4. Results §3.3 / Fig. 6–7: Design and Pnear enrichment are important applications but currently retrospective and qualitative (three campaigns, Kd ≤ 50 µM hits; motif miniaturization examples). To support the claim that Vilya-1 is an “oracle for designing novel topologies,” report prospective or held-out design metrics: recovery of known binders under fixed compute, diversity of accepted scaffolds, and failure modes when the motif is held fixed. Without quantitative design benchmarks, the generative section remains a demonstration rather than evidence for the foundation-model framing.
minor comments (7)
  1. Results §3.2: “combined sampling-and-scoring results shown in Fig 2A” appears to refer to the confidence analysis; Fig. 2 is the architecture schematic. Likely meant Fig. 4A—please correct.
  2. Methods §2.1: Architecture is described at a high level (triangle attention, pair bias, unified diffusion trunk) but omits parameter count, layer depths, embedding sizes, noise schedule, and number of denoising steps. These are free parameters listed implicitly by the work and matter for reproducibility and comparison to Boltz-2/RF3.
  3. Fig. S1 / §2.4: The ring-scaffold definition is a strength; consider promoting a short formal definition into the main text earlier so readers understand why ring RMSD is stricter than Cα-only metrics used elsewhere.
  4. Fig. 1 caption vs. panels B–D: Caption states 100 conformers and no ranking for B–D, while panel A uses top-5 by confidence. Keep this distinction prominent in the main Results text to avoid over-reading sampling-only bars.
  5. §3.1 small-molecule bound results are deferred to Fig. S5; a one-sentence numerical summary in the main text would better support the “extends to small molecules” claim.
  6. Supplement lists CSD codes and PDB IDs (helpful); also list the 66 X-ray test PDB/CSD IDs in one place for independent re-evaluation.
  7. Typos/style: “Schr¨odinger” encoding artifacts; “α-Amanitin” / “SFT1-1” vs “SFTI-1” inconsistency; “treat_bad_torsions fix” formatting.

Circularity Check

0 steps flagged

No circularity: empirical ML model with held-out experimental benchmarks; performance claims do not reduce to training inputs by construction.

full rationale

Vilya-1 is a standard deep-learning conformer generator (all-atom diffusion transformer) trained on mixed public crystal + computational ensembles and evaluated on explicitly held-out X-ray/NMR/receptor-bound macrocycle and small-molecule sets (Methods 2.3–2.4). Success rates (e.g., 89.2 % ring-RMSD < 1 Å on the 66-structure X-ray set) are measured by sampling 100 conformers and taking min RMSD to experimental ground truth; the paper states that molecules >70 % Tanimoto-similar to test cases were excluded from training and that the Fig. 1A examples and the 240-ligand bound set contain no training structures. Confidence and property heads are fine-tuned from the generator weights (standard transfer learning) and ablated against random-init and external baselines (ChemProp/CheMeleon, MLIP, KNN); splits are time-based (internal) or scaffold-based (external PAMPA). Design/Pnear applications use the sampler as an oracle, not as a definitional identity. No equation, loss, or ranking metric is algebraically forced by its own training target; self-citations (e.g., Salveson 2024, Rettie 2025) supply prior design examples or context and are not load-bearing uniqueness theorems. The paper is therefore self-contained against external experimental benchmarks and exhibits none of the six circularity patterns.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 1 invented entities

The central geometric claims rest on standard ML inductive biases plus domain conventions about what constitutes a successful conformer prediction. Free parameters are the usual neural-network and diffusion hyperparameters (not numerically disclosed). No new physical entities are postulated; the model is an empirical function approximator. The main domain assumptions are that experimental crystal/NMR/bound structures are valid targets for ‘biologically relevant’ sampling and that ring-scaffold RMSD < 1 Å is a meaningful success criterion.

free parameters (4)
  • diffusion noise schedule and number of denoising steps
    Standard free choices in diffusion models; exact schedule not reported, yet sampling performance depends on it.
  • model capacity (layers, embedding dimension, attention heads)
    Architecture size is not specified; capacity is a free hyperparameter that affects both accuracy and generalization.
  • loss weights for distogram and chirality auxiliary terms
    Auxiliary losses are introduced without reported coefficients; they influence the final coordinate distribution.
  • ring-RMSD success threshold of 1 Å (and 0.5 Å for small molecules)
    Threshold is conventional but chosen by the authors; success rates are sensitive to it.
axioms (4)
  • domain assumption Experimental X-ray, NMR and receptor-bound coordinates constitute the ground-truth low-energy or biologically relevant conformations that a sampler should recover.
    Stated throughout Results §3.1 and Evaluation Methodology; without this the RMSD success metric loses meaning.
  • ad hoc to paper A purely chemical all-atom feature set (no residue type, atom name or positional encoding) is sufficient for the network to learn transferable structural principles.
    Explicit design choice in §2.1; the paper’s generalization claims rest on it.
  • domain assumption Ring-scaffold RMSD (central ring + direct substituents + fused rings) better captures torsional correctness than Cα-only or main-chain RMSD.
    Defined and justified in §2.4 and Fig. S1; all headline success rates use this definition.
  • domain assumption Time-based splits on internal assay data and scaffold splits on external PAMPA data prevent leakage for property-prediction evaluation.
    Stated in §2.2; enrichment-factor claims depend on the validity of these splits.
invented entities (1)
  • Vilya-1 (the specific all-atom diffusion architecture and its fine-tuned confidence/property heads) no independent evidence
    purpose: Unified model for conformer sampling, ranking, property prediction and discrete design of macrocycles.
    The paper’s contribution is the trained model itself; no independent public weights or formal specification exist outside this work.

pith-pipeline@v1.1.0-grok45 · 23174 in / 3605 out tokens · 36962 ms · 2026-07-14T01:08:36.716873+00:00 · methodology

0 comments
read the original abstract

Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space. In this work, we introduce Vilya-1, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability. Vilya-1 operates on a uniform all-atom representation and is trained on heterogeneous structural datasets spanning diverse topologies and chemical classes. Across a broad set of macrocycles composed of canonical and non-canonical residues, Vilya-1 substantially improves geometric accuracy relative to physics-based methods, co-folding networks, and deep-learning conformer generators, while maintaining broad chemical coverage that extends to small molecules. Vilya-1 also supports generative applications, enabling the design of novel macrocycles with tailored chemical, structural, and property profiles. Together, these capabilities establish Vilya-1 as a foundation model for accelerating the development of next-generation macrocycle therapeutics.

Figures

Figures reproduced from arXiv: 2607.09998 by Adam P. Moyer, Benjamin D. Sellers, Ivan Anishchanka, Milad Salem, Naozumi Hiranuma, Patrick J. Salveson, Stephen Rettie, Vilya Research: Pascal Sturmfels, Xiaoliang Pan.

Figure 1
Figure 1. Figure 1: Bioactive conformer prediction and sampling efficiency of Vilya-1. A) Examples of bioactive macrocycles for which Vilya-1 predicts conformations within sub-angstrom accuracy, measured by ring RMSD. Vilya-1-generated structures (green) are overlaid with the corresponding experimental X-ray structures from the PDB (gray). B) Sampling efficiency of Vilya-1 on X-ray structures of cyclic peptides, compared with… view at source ↗
Figure 2
Figure 2. Figure 2: Schematic of Vilya-1 architecture. Model architecture and noise process components are indicated in green, while feature embeddings and inputs are indicated in gray. The model differs from popular protein ligand co-folding architectures in three main ways. First, the diffusion process runs through a single, unified transformer architecture. As our goal is to explicitly model an ensemble of conformations, w… view at source ↗
Figure 3
Figure 3. Figure 3: Generalization and chemical validity of Vilya-1 on a held-out set of small molecules. A) Vilya-1 accuracy as a function of Tanimoto similarity to the training set, showing no systematic dependence on similarity and indicating that performance is not driven by memorization. B) Chemical validity of Vilya-1-generated conformers assessed using PoseBusters, demonstrating validity rates comparable to reference X… view at source ↗
Figure 4
Figure 4. Figure 4: Conformer scoring with the Vilya-1 confidence model. A) Gain in top-1 success rates (ring RMSD < 1˚A) relative to random selection when conformers are ranked by the Vilya-1 confidence model, MLIP-based energy scoring, and Boltz-2, compared against the empirical upper bound (indicated by the “Lowest RMSD”). The “no transfer” bar corresponds to a Vilya-1 confidence model trained from random initialization ra… view at source ↗
Figure 5
Figure 5. Figure 5: Property prediction with Vilya-1: A) Enrichment factor (top 10%) for property prediction on the internal dataset. All models are evaluated as ensembles of five, with error bars showing 95% confidence intervals. Vilya-1 is initialized from the conformer generator. B) Performance on the hydrophobicity regression task. C) Enrichment factor for permeability prediction on the external benchmark. D) Effect of tr… view at source ↗
Figure 6
Figure 6. Figure 6: A) Vilya-1-generated predictions and B) energy landscapes (upper row) of an experimentally validated heterochiral cyclic peptide (lower row, PDB ID 6BE7) [51] and of a point mutant that replaces D-proline with glycine (lower row). This mutation results in a sequence that stabilizes an entirely different conformation than that adopted by the parent. Side-chains other than proline or D-proline are hidden for… view at source ↗
Figure 7
Figure 7. Figure 7: Computational design and optimization of peptides using non-canonical amino acids and linkages. A) Two thioether macrocycle binders identified from mRNA display, the residues highlighted in green make up the majority of the interface. Using our design methodology, Vilya-1 can be used to miniaturize large peptides using non-canonical amino acids, while maintaining the conformation of the motif. B) From the … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

    cs.LG 2026-07 conditional novelty 6.0

    Vilya-2 predicts bound structures of chemically diverse peptides and small molecules at state-of-the-art accuracy using an all-atom diffusion transformer, recovering 59.1% of peptide interfaces to sub-2 Å backbone RMSD.

Reference graph

Works this paper leans on

61 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Macrocycles in drug discovery - learning from the past for the future

    Diego Garcia Jimenez, Vasanthanathan Poongavanam, and Jan Kihlberg. “Macrocycles in drug discovery - learning from the past for the future”. In:Journal of medicinal chemistry66.8 (2023), pp. 5377–5396

  2. [2]

    Macrocycles for conventionally druggable targets: lessons from macrocyclic kinase inhibitors

    Lauren A Viarengo-Baker and Adrian Whitty. “Macrocycles for conventionally druggable targets: lessons from macrocyclic kinase inhibitors”. In:Journal of Medicinal Chemistry68.15 (2025), pp. 15260–15284

  3. [3]

    Accurate de novo design of hyperstable constrained peptides

    Gaurav Bhardwaj et al. “Accurate de novo design of hyperstable constrained peptides”. In:Nature538.7625 (2016), pp. 329–335

  4. [4]

    Accurate de novo design of high-affinity protein-binding macrocycles using deep learning

    Stephen A Rettie et al. “Accurate de novo design of high-affinity protein-binding macrocycles using deep learning”. In:Nature Chemical Biology(2025), pp. 1–9

  5. [5]

    De novo design of protein structure and function with RFdiffusion

    Joseph L Watson et al. “De novo design of protein structure and function with RFdiffusion”. In:Nature620.7976 (2023), pp. 1089–1100

  6. [6]

    Expansive discovery of chemically diverse structured macrocyclic oligoamides

    Patrick J Salveson et al. “Expansive discovery of chemically diverse structured macrocyclic oligoamides”. In:Science 384.6694 (2024), pp. 420–428

  7. [7]

    John Moult et al.A large-scale experiment to assess protein structure prediction methods. 1995

  8. [8]

    Boltz-2: Towards accurate and efficient binding affinity prediction

    Saro Passaro et al. “Boltz-2: Towards accurate and efficient binding affinity prediction”. In:BioRxiv(2025)

  9. [9]

    Accelerating biomolecular modeling with atomworks and rf3

    Nathaniel Corley et al. “Accelerating biomolecular modeling with atomworks and rf3”. In:bioRxiv(2025)

  10. [10]

    RareFold: Structure prediction and design of proteins with noncanonical amino acids

    Qiuzhen Li et al. “RareFold: Structure prediction and design of proteins with noncanonical amino acids”. In:bioRxiv (2025), pp. 2025–05

  11. [11]

    Accurate structure prediction of cyclic peptides containing unnatural amino acids using HighFold3

    Sen Cao et al. “Accurate structure prediction of cyclic peptides containing unnatural amino acids using HighFold3”. In:Briefings in Bioinformatics26.5 (2025), bbaf488

  12. [12]

    Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2

    Cheng Zhu et al. “Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2”. In: Briefings in Bioinformatics26.3 (2025), bbaf202

  13. [13]

    NCPepFold: Accurate Prediction of Noncanonical Cyclic Peptide Structures via Cyclization Optimization with Multigranular Representation

    Qingyi Mao et al. “NCPepFold: Accurate Prediction of Noncanonical Cyclic Peptide Structures via Cyclization Optimization with Multigranular Representation”. In:Journal of Chemical Theory and Computation21.9 (2025), pp. 4979–4991

  14. [14]

    Improving accuracy, diversity, and speed with prime macrocycle conformational sampling

    Dan Sindhikara et al. “Improving accuracy, diversity, and speed with prime macrocycle conformational sampling”. In:Journal of chemical information and modeling57.8 (2017), pp. 1881–1894

  15. [15]

    Complex macrocycle exploration: parallel, heuristic, and constraint-based conformer generation using ForceGen

    Ajay N Jain et al. “Complex macrocycle exploration: parallel, heuristic, and constraint-based conformer generation using ForceGen”. In:Journal of computer-aided molecular design33.6 (2019), pp. 531–558. Vilya Research 13

  16. [16]

    Elucidating solution structures of cyclic peptides using molecular dynamics simulations

    Jovan Damjanovic et al. “Elucidating solution structures of cyclic peptides using molecular dynamics simulations”. In:Chemical reviews121.4 (2021), pp. 2292–2324

  17. [17]

    Accurate structure prediction of biomolecular interactions with AlphaFold 3

    Josh Abramson et al. “Accurate structure prediction of biomolecular interactions with AlphaFold 3”. In:Nature 630.8016 (2024), pp. 493–500

  18. [18]

    Modeling protein–small molecule conformational ensembles with PLACER

    Ivan Anishchenko et al. “Modeling protein–small molecule conformational ensembles with PLACER”. In:Proceedings of the National Academy of Sciences122.45 (2025), e2427161122

  19. [19]

    Proteina: Scaling flow-based protein structure generative models

    Tomas Geffner et al. “Proteina: Scaling flow-based protein structure generative models”. In:arXiv preprint arXiv:2503.00710(2025)

  20. [20]

    Highly accurate protein structure prediction with AlphaFold

    John Jumper et al. “Highly accurate protein structure prediction with AlphaFold”. In:nature596.7873 (2021), pp. 583–589

  21. [21]

    Generalized biomolecular modeling and design with RoseTTAFold All-Atom

    Rohith Krishna et al. “Generalized biomolecular modeling and design with RoseTTAFold All-Atom”. In:Science 384.6693 (2024), eadl2528

  22. [22]

    CycPeptMPDB: a comprehensive database of membrane permeability of cyclic peptides

    Jianan Li et al. “CycPeptMPDB: a comprehensive database of membrane permeability of cyclic peptides”. In: Journal of Chemical Information and Modeling63.7 (2023), pp. 2240–2250

  23. [23]

    Qiushi Feng.Development of an Open-source Non-peptidic Macrocycle Membrane Permeability Database and Predictive Machine Learning Models. 2024

  24. [24]

    MolMeDB: molecules on membranes database

    Jakub Jura ˇcka et al. “MolMeDB: molecules on membranes database”. In:Database2019 (2019), baz078

  25. [25]

    The NCATS Pharmaceutical Collection: a 10-year update

    Ruili Huang et al. “The NCATS Pharmaceutical Collection: a 10-year update”. In:Drug Discovery Today24.12 (2019), pp. 2341–2349

  26. [26]

    PerMM: a web tool and database for analysis of passive membrane permeability and translocation pathways of bioactive molecules

    Andrei L Lomize et al. “PerMM: a web tool and database for analysis of passive membrane permeability and translocation pathways of bioactive molecules”. In:Journal of chemical information and modeling59.7 (2019), pp. 3094–3099

  27. [27]

    Conformational effects on the passive membrane permeability of synthetic macrocycles

    Anna A Rzepiela et al. “Conformational effects on the passive membrane permeability of synthetic macrocycles”. In:Journal of medicinal chemistry65.15 (2022), pp. 10300–10317

  28. [28]

    SIMPD: an algorithm for generating simulated time splits for validating machine learning approaches

    Gregory A Landrum et al. “SIMPD: an algorithm for generating simulated time splits for validating machine learning approaches”. In:Journal of cheminformatics15.1 (2023), p. 119

  29. [29]

    The Cambridge structural database

    Colin R Groom et al. “The Cambridge structural database”. In:Structural Science72.2 (2016), pp. 171–179

  30. [30]

    Accurate de novo design of membrane-traversing macrocycles

    Gaurav Bhardwaj et al. “Accurate de novo design of membrane-traversing macrocycles”. In:Cell185.19 (2022), pp. 3520–3532

  31. [31]

    Cyclic peptide structure prediction and design using AlphaFold2

    Stephen A Rettie et al. “Cyclic peptide structure prediction and design using AlphaFold2”. In:Nature Communications 16.1 (2025), p. 4730

  32. [32]

    Biotite: a unifying open source computational biology framework in Python

    Patrick Kunzmann and Kay Hamacher. “Biotite: a unifying open source computational biology framework in Python”. In:BMC bioinformatics19.1 (2018), p. 346

  33. [33]

    Version Release 2025 09 4

    Greg Landrum.rdkit/rdkit: 2025 09 4 (Q3 2025) Release. Version Release 2025 09 4. Dec. 2025.doi: 10.5281/ zenodo.18098214.url:https://doi.org/10.5281/zenodo.18098214

  34. [34]

    High-quality dataset of protein-bound ligand conformations and its application to benchmarking conformer ensemble generators

    Nils-Ole Friedrich et al. “High-quality dataset of protein-bound ligand conformations and its application to benchmarking conformer ensemble generators”. In:Journal of chemical information and modeling57.3 (2017), pp. 529–539

  35. [35]

    Accurate physics-based flexible docking of macrocyclic ligands

    Jacob Robson-Tull and Jo˜ao Rodrigues. “Accurate physics-based flexible docking of macrocyclic ligands”. In: (2025)

  36. [36]

    Improving conformer generation for small rings and macrocycles based on distance geometry and experimental torsional-angle preferences

    Shuzhe Wang et al. “Improving conformer generation for small rings and macrocycles based on distance geometry and experimental torsional-angle preferences”. In:Journal of chemical information and modeling60.4 (2020), pp. 2044–2058

  37. [37]

    Scalable Low-Energy Molecular Conformer Generation with Quantum Mechanical Accuracy

    Filipp Nikitin et al. “Scalable Low-Energy Molecular Conformer Generation with Quantum Mechanical Accuracy”. In: (2025)

  38. [38]

    Torsional diffusion for molecular conformer generation

    Bowen Jing et al. “Torsional diffusion for molecular conformer generation”. In:Advances in neural information processing systems35 (2022), pp. 24240–24253

  39. [39]

    Et-flow: Equivariant flow-matching for molecular conformer generation

    Majdi Hassan et al. “Et-flow: Equivariant flow-matching for molecular conformer generation”. In:Advances in Neural Information Processing Systems37 (2024), pp. 128798–128824. Vilya Research 14

  40. [40]

    Grambow et al.Accurate and Efficient Structural Ensemble Generation of Macrocyclic Peptides using Internal Coordinate Diffusion

    Colin A. Grambow et al.Accurate and Efficient Structural Ensemble Generation of Macrocyclic Peptides using Internal Coordinate Diffusion. 2024. arXiv:2305.19800 [q-bio.BM].url:https://arxiv.org/abs/2305.19800

  41. [41]

    Learning smooth and expressive interatomic potentials for physical property prediction

    Xiang Fu et al. “Learning smooth and expressive interatomic potentials for physical property prediction”. In:arXiv preprint arXiv:2502.12147(2025)

  42. [42]

    Descriptor-based Foundation Models for Molecular Property Prediction

    Jackson Burns, Akshat Zalte, and William Green. “Descriptor-based Foundation Models for Molecular Property Prediction”. In:arXiv preprint arXiv:2506.15792(2025)

  43. [43]

    Chemprop: a machine learning package for chemical property prediction

    Esther Heid et al. “Chemprop: a machine learning package for chemical property prediction”. In:Journal of Chemical Information and Modeling64.1 (2023), pp. 9–17

  44. [44]

    May 2025.doi: 10

    Jackson Burns.CheMeleon Foundation Model. May 2025.doi: 10 . 5281 / zenodo . 15426601.url: https : //doi.org/10.5281/zenodo.15426601

  45. [45]

    Have protein-ligand co-folding methods moved beyond memorisation?

    Peter ˇSkrinjar et al. “Have protein-ligand co-folding methods moved beyond memorisation?” In:BioRxiv(2025), pp. 2025–02

  46. [46]

    PLINDER: The protein-ligand interactions dataset and evaluation resource

    Janani Durairaj et al. “PLINDER: The protein-ligand interactions dataset and evaluation resource”. In:BioRxiv (2024), pp. 2024–07

  47. [47]

    Deep learning for protein-ligand docking: Are we there yet?

    Alex Morehead et al. “Deep learning for protein-ligand docking: Are we there yet?” In:ArXiv(2025), arXiv–2405

  48. [48]

    PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences

    Martin Buttenschoen, Garrett M Morris, and Charlotte M Deane. “PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences”. In:Chemical Science15.9 (2024), pp. 3130–3139

  49. [49]

    AIMNet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs

    Dylan M Anstine, Roman Zubatyuk, and Olexandr Isayev. “AIMNet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs”. In:Chemical Science16.23 (2025), pp. 10228–10244

  50. [50]

    The open molecules 2025 (omol25) dataset, evaluations, and models

    Daniel S Levine et al. “The open molecules 2025 (omol25) dataset, evaluations, and models”. In:arXiv preprint arXiv:2505.08762(2025)

  51. [51]

    Comprehensive computational design of ordered peptide macrocycles

    Parisa Hosseinzadeh et al. “Comprehensive computational design of ordered peptide macrocycles”. In:Science 358.6369 (2017), pp. 1461–1466

  52. [52]

    State-selective modulation of heterotrimeric G αs signaling with macrocyclic peptides

    Shizhong A Dai et al. “State-selective modulation of heterotrimeric G αs signaling with macrocyclic peptides”. In: Cell185.21 (2022), pp. 3950–3965

  53. [53]

    In vitro selection of macrocyclic peptide inhibitors containing cyclic γ2, 4-amino acids targeting the SARS-CoV-2 main protease

    Takashi Miura et al. “In vitro selection of macrocyclic peptide inhibitors containing cyclic γ2, 4-amino acids targeting the SARS-CoV-2 main protease”. In:Nature Chemistry15.7 (2023), pp. 998–1005

  54. [54]

    Approaches for peptide and protein cyclisation

    Heather C Hayes, Louis YP Luk, and Yu-Hsuan Tsai. “Approaches for peptide and protein cyclisation”. In:Organic & biomolecular chemistry19.18 (2021), pp. 3983–4001

  55. [55]

    mRNA display: from basic principles to macrocycle drug discovery

    Kristopher Josephson, Alonso Ricardo, and Jack W Szostak. “mRNA display: from basic principles to macrocycle drug discovery”. In:Drug Discovery Today19.4 (2014), pp. 388–399

  56. [56]

    RNA display methods for the discovery of bioactive macrocycles

    Yichao Huang, Mareike Margarete Wiedmann, and Hiroaki Suga. “RNA display methods for the discovery of bioactive macrocycles”. In:Chemical reviews119.17 (2018), pp. 10360–10391

  57. [57]

    Ultra-large chemical libraries for the discovery of high-affinity peptide binders

    Anthony J Quartararo et al. “Ultra-large chemical libraries for the discovery of high-affinity peptide binders”. In: Nature communications11.1 (2020), p. 3183

  58. [58]

    The RaPID platform for the discovery of pseudo-natural macrocyclic peptides

    Yuki Goto and Hiroaki Suga. “The RaPID platform for the discovery of pseudo-natural macrocyclic peptides”. In: Accounts of Chemical Research54.18 (2021), pp. 3604–3617

  59. [59]

    Macrocyclic DNA-encoded chemical libraries: a historical perspective

    Louise Plais and J ¨org Scheuermann. “Macrocyclic DNA-encoded chemical libraries: a historical perspective”. In: RSC Chemical Biology3.1 (2022), pp. 7–17

  60. [60]

    Validation of a new methodology to create oral drugs beyond the rule of 5 for intracellular tough targets

    Atsushi Ohta et al. “Validation of a new methodology to create oral drugs beyond the rule of 5 for intracellular tough targets”. In:Journal of the American Chemical Society145.44 (2023), pp. 24035–24051

  61. [61]

    De novo mapping of α-helix recognition sites on protein surfaces using unbiased libraries

    Kunhua Li et al. “De novo mapping of α-helix recognition sites on protein surfaces using unbiased libraries”. In: Proceedings of the National Academy of Sciences119.52 (2022), e2210435119. Vilya Research 15 A Supplement A.1 Macrocycle validation set AAGAGG10, ABENEC, ADUGAH, ADUHUC, AFOKUE, AFUWON, AMEHUX, ANUJUQ, ANUTOU, AXAWAZ, BAPLUD, BAXBUB, BAXPEX, B...