Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Latent-X jointly designs protein binder sequence and all-atom structure, reporting 91-100% macrocycle hit rates and picomolar affinities in wet-lab tests.

desk verdict Real wet-lab results and a credible 200-target benchmark, but the headline macrocycle hit rates rest on small post-filter subsets and the model is proprietary, so the 'as few as 30–100 designs' claim overstates what is actually demonstrated. read the letter →

arxiv 2507.19375 v1 pith:CVFWRPHQ submitted 2025-07-25 q-bio.BM

classification q-bio.BM
keywords denovoproteinbinderdesignall-atomgenerativemodelmacrocyclicpeptidesmini-bindersepitopeconditioningstructurepredictionfilteringwet-labvalidationbindingaffinity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces Latent-X, a generative model that takes a target protein and a few hotspot residues and produces both the amino-acid sequence and the all-atom structure of a protein binder, while also co-generating the target side chains at the interface. The central claim is that binder design can be reduced to testing a few dozen to a hundred designs per target rather than screening millions of molecules. In wet-lab experiments across seven targets the paper reports macrocycle hit rates from 91% to 100%, mini-binder hit rates from 10% to 64%, and affinities reaching below 0.01 nanomolar, with the best numbers beating previously reported binders from the strongest existing methods under identical assay conditions. The paper further claims that the model runs an order of magnitude faster than multi-step pipelines and produces structurally diverse binders, including beta-sheet folds.

What carries the argument

The load-bearing mechanism is joint co-generation of sequence and all-atom structure for binder and target in a single generative pass, in contrast to conventional pipelines that first sample a backbone and then assign a sequence with a separate model; this lets the model optimise interface side-chain rotamers and hydrogen-bonding networks directly. Around the generator sits an in silico filter built on structure-prediction models, using thresholds on the minimum interface predicted aligned error (min_ipae), the binder's predicted TM-score (ptm_binder) and the RMSD between generated and predicted complex (complex_rmsd), which selects the handful of designs that go to the lab.

What would settle it

Test the claim by taking a panel of, say, 20 target proteins unrelated to the filter-tuning set, running the complete Latent-X workflow, and synthesising a random sample of both pass and fail designs from the in silico filter; the paper's claim predicts a large gap in experimental hit rates between pass and fail sets, while the filter-bias objection predicts comparable rates, especially on non-helical folds. A simpler check: for one macrocycle target, synthesise the roughly 600 designs that did not pass the filter and measure what fraction still bind, since a >90% generation-level hit rate predicts many should.

Watch

Extended reading notes

Core claim

Latent-X is an all-atom generative model that simultaneously predicts the sequence and full atomic structure of both the binder and the target protein in one pass. Prompted with a target structure and hotspot residues, it directly constructs non-covalent interactions, including extensive interface hydrogen-bond networks, instead of leaving interface chemistry to post-hoc refinement. In the authors' experimental validation, macrocycles against MDM2, MCL-1 and PD-L1 achieved hit rates of 90.9%, 100% and 94.1% respectively with best affinities of 5.35 micromolar, 18.4 micromolar and 71.7 micromolar, and mini-binders against BHRF1, TrkA, PD-L1, IL-7R-alpha and SC2RBD achieved hit rates of 64%, 10%, 49%, 26% and 52% with best affinities ranging from 32.4 nanomolar on BHRF1 to below 0.01 nanomolar on IL-7R-alpha and SC2RBD. These results are presented as outperforming the best published binders from RFdiffusion, RFpeptides and AlphaProteo when replicated in identical assays, with an order-of-magnitude faster generation and an in silico hit-rate advantage on 200 held-out PDB targets.

Load-bearing premise

The load-bearing premise is that the in silico filters built on structure-prediction confidence metrics pick out designs that will actually bind in the lab; the paper itself concedes these filters were tuned on helical 45-65 residue mini-binders and that the macrocycle filter had to be relaxed after known binders failed it, so if the filters select for easy-to-predict rather than genuinely potent proteins, the reported hit rates overstate the model's stand-alone generation quality.

Editorial extensions

If this is right

  • If the wet-lab numbers hold, a researcher can go from a target structure to candidates worth testing with 30-100 designs per target instead of millions of screened molecules.
  • Head-to-head superiority under identical conditions would make Latent-X the default generative method for macrocyclic peptide and mini-binder design, displacing backbone-diffusion-plus-resequencing workflows.
  • An order-of-magnitude inference speed-up makes design interactive and allows large in silico sweeps over targets and epitopes as a routine step.
  • Structural diversity, including beta-sheet folds, suggests the accessible design space extends beyond the helical bundles that previous methods mainly produced.
  • Zero-shot generalisation to both cyclic and linear topologies without task-specific fine-tuning implies one model can serve multiple therapeutic modalities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the authors leave implicit: the reported macrocycle hit rates are for the full pipeline of generator plus relaxed macrocycle filter, so the generator's standalone success rate on unfiltered designs is probably lower; the paper's own in silico success rates of 56-67% before synthesis filters hint at the magnitude.
  • The disclosed filter bias toward 45-65 residue helical mini-binders suggests the 10% hit rate on TrkA and 26% on IL-7R-alpha may understate what the model itself achieves on non-helical folds; testing unfiltered or filter-agnostic sets would quantify this.
  • A testable extension: because conditioning is via hotspot residues, the same pipeline could be asked to design binders with prescribed multi-target specificity or to hit two epitopes on one target, which the paper only gestures at in future work.
  • The in silico benchmark on 200 held-out PDB targets predicts that some fraction of arbitrary targets, roughly 5.5% for mini-binders by the paper's own count, produce no passing designs; a prospective user should expect occasional total failures and plan around them.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. Latent-X is an all-atom generative model that, given a target structure and hotspot residues, co-generates the structure and sequence of a protein binder together with the bound target conformation. The authors validate the full pipeline (generation plus in silico filtering) in wet-lab experiments across two modalities: macrocyclic peptides (12–18 residues) against MDM2, MCL-1, and PD-L1, and mini-binders (80–120 residues) against BHRF1, IL-7Ra, PD-L1, SC2RBD, and TrkA. Headline claims are experimental hit rates of 91–100% for macrocycles (computed on 11–17 SPR-tested designs per target) and 10–64% for mini-binders (100 HT-BLI-tested per target, with reported KD values down to <0.01 nM for several designs), specificity in an orthogonal mammalian-display assay, an order-of-magnitude inference speed advantage over RFdiffusion, and higher expected in silico hit rates than RFpeptides/RFdiffusion on 200 held-out PDB targets (8.26% vs 1.72% for macrocycles; 5.11% vs 3.02% for mini-binders). The paper also reports replicated affinity measurements of previously published AlphaProteo, RFdiffusion, and RFpeptides binders under the same assay conditions.

Significance. If the mini-binder results are taken as reported, they represent a genuinely strong empirical contribution: 100 designs per target were tested with a stated response threshold (0.03 RU), hit rates range from 10% to 64%, affinities were confirmed by 5-point BLI on a subset, and an orthogonal mDisplay assay correlates with HT-BLI (Pearson r = 0.68–0.79). The prospective in silico benchmark on 200 structures deposited after the training cutoff of both the design and structure-prediction models, using openly available Chai-1/Boltz-2 metrics and repeated for both filters, is a serious attempt at estimating generalization, and the paper provides the data needed to check the reported funnel. Credit is also due for the explicit disclosure of the post-hoc macrocycle filter relaxation (App. B.1), the mini-binder filter bias (Sec. 5.3), and synthesis attrition (Sec. 3.1), and for reporting the weaker macrocycle KD values in Fig. S8 rather than only the best ones.

major comments (3)
  1. [Sec. 3.1, Table 2, Fig. S8; Abstract] The headline macrocycle claim is not supported by the reported data. Section 3.1 states that 30 designs per target were submitted for synthesis, that only 57% (MDM2), 77% (MCL-1), and 87% (PD-L1) cyclized at >90% purity, and that 11–17 designs per target were selected for SPR. The hit rates in Table 2 (90.9% for MDM2, 100% for MCL-1, 94.1% for PD-L1) are therefore 10/11, 11/11, and 16/17, computed on a small, non-random subset of the submitted designs. The selection criteria for this subset ('chosen to provide a representative sample') are not specified, so the abstract's 'testing as few as 30–100 designs per target' and 'hit rates exceeding 90%' are misleading; even under the optimistic assumption that every untested cyclized design binds, the per-submitted-design ceilings are 33–53%. Second, the hit definition ('measurable binding', Table 2 footnote) counts SPR designs whose fitted KD is in the millimolar range, several orders of magnitude above the highest analyte concentration used (25–500 µM, Fig. S12), for example LL_CYC_PD-L1_22 at 10.4 mM and LL_CYC_MCL-1_16 at 1.34 mM in Fig. S8; these are extrapolated fits, not established binding. Applying a KD < 100 µM threshold to the values in Fig. S8 reduces the hit rates to roughly 5/11 for MDM2, 6–7/11 for MCL-1, and 3–4/17 for PD-L1. Please report the full funnel (generated, filter-passed, submitted, cyclized, tested, bound at an explicit threshold) and revise the abstract and Sec. 5.1 accordingly.
  2. [Abstract, Table 2, Sec. 5.2] The claim of 'direct comparisons ... under identical conditions ... higher hit rates' (abstract) is contradicted by the paper's own table. Table 2's footnote states that AlphaProteo and RFdiffusion/RFpeptides hit rates are 'taken from the original publications,' which used different assays and thresholds (e.g., AlphaProteo's yeast-display pipeline versus the present HT-BLI); only the affinity measurements in Table 1, on replicated binder sequences, were made under identical conditions. Furthermore, the paper's own Table 2 shows AlphaProteo's published hit rate of 88% on BHRF1 exceeding Latent-X's 64%, and RFdiffusion's published 33.7% on IL-7Ra exceeding Latent-X's 26.0%. The blanket statements in the abstract and in Sec. 5.2 ('consistent advantages across all tested targets') therefore overstate the evidence. Please either restrict the hit-rate claims to the targets and assays where they hold, or report measured hit rates for the comparator workflows under the same assay and threshold.
  3. [App. B.1, Sec. 3.1] The macrocycle in silico filter was modified after observing failures. App. B.1 reports that the ptm_binder criterion was dropped because three known experimentally validated macrocycle complexes (9cdz, 7oun, 1sfi) failed it, and that 'this choice of filter for macrocycles was post-hoc justified by the high resulting experimental success.' Because the 11–17 SPR-tested macrocycles per target were selected from the top-30 of 700 by this relaxed filter, the >90% hit rates are conditional on a filter whose key modification was motivated by the outcome the filter is then used to validate. The paper should (i) report how many of the 11–17 tested macrocycles would have passed the unmodified filter, (ii) present the relaxation as a limitation at the level of the abstract rather than only in an appendix, and (iii) attribute the reported rate to the full model-plus-filter selection pipeline rather than to the generative model alone.
minor comments (6)
  1. [Fig. S8] The caption states '25 de novo designed bound structures across 3 different targets,' but the figure lists 37 structures (10 MDM2, 11 MCL-1, 16 PD-L1); the caption should be corrected.
  2. [Sec. 3.2, Tab. S4] KD values reported as '<0.01 nM' (e.g., LL_MINI_IL-7Ra_39, LL_MINI_SC2RBD_8) are below the reliable quantification range of the 5-point BLI assay and should be reported as upper bounds (KD < 10 pM) to avoid overstating precision.
  3. [Abstract] The sentence beginning 'In direct comparisons with the state-of-the-art models ... under identical conditions demonstrates' is grammatically incomplete and should be rewritten.
  4. [Sec. 2.1 vs Sec. 4] The experimental results are from Latent-X v1, whereas the in silico expected hit-rate benchmark (Fig. 6) uses v1.1; this version mismatch should be stated wherever Fig. 6 results are cited, since the reader cannot judge whether the two sets of numbers describe the same model.
  5. [Sec. 3.1, Fig. 5, Tab. S3] The target name is rendered as 'IL-7Rα' in some places and 'IL-7Ra' in others; the notation should be unified.
  6. [Sec. 3.2] The number of Tier-2 designs per target advanced from HT-BLI to 5-point BLI, and the precise selection rule (thresholds on Ka and RU), are not reported; without these, the KD distribution in Fig. S9 cannot be related to the HT-BLI hit rates.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Latent-X's claims rest on external wet-lab measurements; the disclosed post-hoc macrocycle filter adjustment and denominator reporting affect generalizability, not derivation.

full rationale

The paper's core quantitative claims are wet-lab hit rates and KD values obtained by SPR, BLI, and mDisplay on designs generated by Latent-X, with published binders replicated in the same assays as controls (Tabs 1-2, App. E). Those measurements are external to the generative model, so the abstract's >90% macrocycle and picomolar mini-binder claims are not the model's outputs re-labeled as predictions. The in silico filter is tuned on the independent Cao et al. dataset (640,000 experimentally characterized complexes, App. B) using thresholds reported in Tab. S1; nothing in that tuning takes Latent-X's own experimental outcomes as input. The disclosed limitation in Sec. 5.3 ('The in silico filters used for design selection were optimized based on lab-validated mini-binders of 45-65 amino acids with predominantly helical folds. This bias may introduce selection artifacts') and the post-hoc relaxation in App. B.1 ('we opted to drop this component in the filter when filtering macrocyle binders... post-hoc justified by the high resulting experimental success') do reduce the strength of the claim that the full pre-specified pipeline achieves the reported hit rates, and the denominator discrepancy (30 submitted vs 11-17 SPR-tested macrocycles, Sec. 3.1 and Tab. 2) limits extrapolation to 'as few as 30 designs'. These are statistical/reporting concerns, not definitional circularity: the reported binders were still selected before wet-lab measurement and their binding was measured independently. There is no self-citation chain, no ansatz smuggled via citation, and no uniqueness theorem imported from the authors. Verdict: no significant circularity (score 0).

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claims rest on an empirical model whose architecture is undisclosed and on a filtering funnel whose thresholds were tuned on external data (Cao et al.) and, for macrocycles, adjusted post-hoc. No new physical entities are introduced. The key unstated inputs are the filter thresholds, the hit-rate definition, and the assumption that structure-prediction confidence correlates with binding.

free parameters (4)
  • Chai-1 in silico filter thresholds = min_ipae < 1, ptm_binder > 0.9, complex_rmsd < 2
    Tuned by grid search over the Cao et al. dataset to maximize precision@1% (App B, Tab. S1). These thresholds determine which designs are selected for lab testing and therefore set the reported hit rates and affinities.
  • Boltz-2 in silico filter thresholds = min_ipae < 1, ptm_binder > 0.95, complex_rmsd < 2.5
    Alternative filter used to confirm in silico hit-rate comparisons (App D.3, Tab. S1).
  • HT-BLI hit threshold = 0.03 RU response
    Designs with HT-BLI response above 0.03 RU after reference subtraction were classified as binders; this directly defines the mini-binder hit rates (App E.5).
  • Macrocycle filter relaxation = drop ptm_binder criterion
    The ptm_binder component was removed from the filter for macrocycles after three known macrocycle binders failed it, changing which macrocycles were lab-tested (App B.1).
assumptions (5)
  • domain assumption The provided crystal structure of the target represents the binding-competent conformation
    The model conditions on a single static structure and the paper lists target-structure availability as a limitation (Sec. 5.3, App A.2).
  • domain assumption Structure-prediction confidence metrics from Chai-1 and Boltz-2 correlate with experimental binding
    The design funnel and the in silico benchmark treat these metrics as binding proxies; the filters were tuned on the Cao et al. dataset (App B).
  • domain assumption Hotspot residues supplied by the user define a targetable binding site
    All designs are conditioned on hand-picked hotspot residues from prior literature or natural partners (Tab. S3); changing hotspots would change results.
  • domain assumption The Cao et al. yeast-display dataset is a valid external standard for tuning filter thresholds
    Thresholds are optimized on this dataset and then applied to macrocycles and non-helical folds (App B).
  • ad hoc to paper Macrocycle designs should be filtered without the ptm_binder criterion
    Added after three known macrocycle binders failed the mini-binder filter; this choice raised the reported macrocycle success rates (App B.1, Tab. S2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design." pith.science (2026). https://pith.science/paper/CVFWRPHQ

@misc{pith2026250719375,
  author       = {Pith},
  title        = {Pith review of: Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CVFWRPHQ}},
  note         = {Machine review of arXiv:2507.19375}
}
read the original abstract

Traditional drug discovery relies on rounds of screening millions of candidate molecules with low success rates, making drug discovery time and resource intensive. To overcome this screening bottleneck, we introduce Latent-X, an all-atom protein design model that enables a new paradigm of precision AI design. Given a target protein epitope, Latent-X jointly generates the all atom structure and sequence of the protein binder and target, directly modelling the non-covalent interactions essential for specific binding. We demonstrate its efficacy across two therapeutically relevant modalities through extensive wet lab experiments, testing as few as 30-100 designs per target. For macrocyclic peptides, Latent-X achieves experimental hit rates exceeding 90% on all evaluated benchmark targets. For mini-binders, it consistently produces potent candidates against all evaluated benchmark targets, with binding affinities reaching the low nanomolar and picomolar range - comparable to those of approved therapeutics - whilst also being highly specific in mammalian display. In direct comparisons with the state-of-the-art models AlphaProteo, RFdiffusion and RFpeptides under identical conditions demonstrates, Latent-X generates binders with higher hit rates and better binding affinities, and uniquely creates structurally diverse binders, including complex beta-sheet folds. Its end-to-end process is an order of magnitude faster than existing multi-step computational pipelines. By drastically improving the efficiency and success rate of de novo design, Latent-X represents a significant advance towards push-button biologics discovery and a valuable tool for protein engineers. Latent-X is available at https://platform.latentlabs.com, enabling users to reliably generate de novo binders without AI infrastructure or coding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Design-CP: Context Parallelism for Design of Protein Nanoparticles

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Context-parallel inference for RFdiffusion 3 enables end-to-end all-atom design of large symmetric protein nanoparticles on multi-GPU hardware without retraining.

Reference graph

Works this paper leans on

61 extracted references · 49 canonical work pages · cited by 1 Pith paper

  1. [1]

    Therapeutic proteins.Therapeutic Proteins: Methods and Protocols, pages 1–26, 2012

    Dimiter S Dimitrov. Therapeutic proteins.Therapeutic Proteins: Methods and Protocols, pages 1–26, 2012

  2. [2]

    Recent advances in the development of protein–protein interactions modulators: mechanisms and clinical trials

    Haiying Lu, Qiaodan Zhou, Jun He, Zhongliang Jiang, Cheng Peng, Rongsheng Tong, and Jianyou Shi. Recent advances in the development of protein–protein interactions modulators: mechanisms and clinical trials. Signal transduction and targeted therapy, 5(1):213, 2020

  3. [3]

    Engineering protein-based therapeutics through structural and chemical design.Nature communications, 14(1):2411, 2023

    Sasha B Ebrahimi and Devleena Samanta. Engineering protein-based therapeutics through structural and chemical design.Nature communications, 14(1):2411, 2023

  4. [4]

    The exploration of macrocycles for drug discovery—an underexploited structural class.Nature Reviews Drug Discovery, 7(7):608–624, 2008

    Edward M Driggers, Stephen P Hale, Jinbo Lee, and Nicholas K Terrett. The exploration of macrocycles for drug discovery—an underexploited structural class.Nature Reviews Drug Discovery, 7(7):608–624, 2008

  5. [5]

    Macrocyclic peptides as drug candidates: Recent progress and remaining challenges.Journal of the American Chemical Society, 141(10):4167–4181, 2019

    Alexander A Vinogradov, Yizhen Yin, and Hiroaki Suga. Macrocyclic peptides as drug candidates: Recent progress and remaining challenges.Journal of the American Chemical Society, 141(10):4167–4181, 2019

  6. [6]

    Massively parallel de novo protein design for targeted therapeutics.Nature, 550(7674):74–79, 2017

    AaronChevalier,Daniel-AdrianoSilva,GabrielJRocklin,DerrickRHicks,RenanVergara,PatienceMurapa, Steffen M Bernard, Lu Zhang, Kwok-Ho Lam, Guorui Yao, et al. Massively parallel de novo protein design for targeted therapeutics.Nature, 550(7674):74–79, 2017

  7. [7]

    De novo design of picomolar SARS-CoV-2 miniprotein inhibitors

    Longxing Cao, Inna Goreshnik, Brian Coventry, James Brett Case, Lauren Miller, Lisa Kozodoy, Rita E Chen, Lauren Carter, Alexandra C Walls, Young-Jun Park, et al. De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science, 370(6515):426–431, 2020

  8. [8]

    An enumerative algorithm for de novo design of proteins with diverse pocket structures.Proceedings of the National Academy of Sciences, 117(36):22135–22145, 2020

    BenjaminBasanta,MatthewJBick,AsimKBera,ChristofferNorn,CameronMChow,LaurenPCarter,Inna Goreshnik, Frank Dimaio, and David Baker. An enumerative algorithm for de novo design of proteins with diverse pocket structures.Proceedings of the National Academy of Sciences, 117(36):22135–22145, 2020

Show all 61 references
  1. [9]

    Design of protein-binding proteins from the target structure alone.Nature, 605(7910):551–560, 2022

    Longxing Cao, Brian Coventry, Inna Goreshnik, Buwei Huang, William Sheffler, Joon Sung Park, Kevin M Jude, Iva Marković, Rameshwar U Kadam, Koen HG Verschueren, et al. Design of protein-binding proteins from the target structure alone.Nature, 605(7910):551–560, 2022

  2. [10]

    Highly accurate protein structure prediction with AlphaFold.nature, 596(7873):583–589, 2021

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al. Highly accurate protein structure prediction with AlphaFold.nature, 596(7873):583–589, 2021. Latent-X: An A...

  3. [11]

    Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem

    Brian L Trippe, Jason Yim, Doug Tischer, David Baker, Tamara Broderick, Regina Barzilay, and Tommi Jaakkola. Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem. arXiv preprint arXiv:2206.04119, 2022

  4. [12]

    Protein structure and sequence generation with equivariant denoising diffusion probabilistic models.arXiv preprint arXiv:2205.15019, 2022

    Namrata Anand and Tudor Achim. Protein structure and sequence generation with equivariant denoising diffusion probabilistic models.arXiv preprint arXiv:2205.15019, 2022

  5. [13]

    Scaffolding protein functional sites using deep learning

    JueWang,SidneyLisanza,DavidJuergens,DougTischer,JosephLWatson,KarlaMCastro,RobertRagotte, Amijai Saragovi, Lukas F Milles, Minkyung Baek, et al. Scaffolding protein functional sites using deep learning. Science, 377(6604):387–394, 2022

  6. [14]

    Illuminating protein space with a programmable generative model.Nature, 623(7989):1070–1078, 2023

    JohnBIngraham,MaxBaranov,ZakCostello,KarlWBarber,WujieWang,AhmedIsmail,VincentFrappier, Dana M Lord, Christopher Ng-Thow-Hing, Erik R Van Vlack, et al. Illuminating protein space with a programmable generative model.Nature, 623(7989):1070–1078, 2023

  7. [15]

    Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design.arXiv preprint arXiv:2402.04997, 2024

    Andrew Campbell, Jason Yim, Regina Barzilay, Tom Rainforth, and Tommi Jaakkola. Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design.arXiv preprint arXiv:2402.04997, 2024

  8. [16]

    Simulating 500 million years of evolution with a language model.Science, page eads0018, 2025

    Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model.Science, page eads0018, 2025

  9. [17]

    An all-atomproteingenerativemodel

    AlexanderEChu,JinhoKim,LucyCheng,GinaElNesr,MinkaiXu,RichardWShuai,andPo-SsuHuang. An all-atomproteingenerativemodel. Proceedings of the National Academy of Sciences,121(27):e2311500121, 2024

  10. [18]

    P(all-atom) is unlocking new path for protein design.bioRxiv, pages 2024–08, 2024

    Wei Qu, Jiawei Guan, Rui Ma, Ke Zhai, Weikun Wu, and Haobo Wang. P(all-atom) is unlocking new path for protein design.bioRxiv, pages 2024–08, 2024

  11. [19]

    All-atom protein generation with latent diffusion

    AmyXLu,WilsonYan,SarahARobinson,SimonKelow,KevinKYang,VladimirGligorijevic,Kyunghyun Cho, Richard Bonneau, Pieter Abbeel, and Nathan C Frey. All-atom protein generation with latent diffusion. In ICLR 2025 Workshop on Generative and Experimental Perspectives for Biomolecular Des...

  12. [20]

    De novo protein design by deep network hallucination.Nature, 600(7889):547–552, 2021

    Ivan Anishchenko, Samuel J Pellock, Tamuka M Chidyausiku, Theresa A Ramelot, Sergey Ovchinnikov, Jingzhou Hao, Khushboo Bafna, Christoffer Norn, Alex Kang, Asim K Bera, et al. De novo protein design by deep network hallucination.Nature, 600(7889):547–552, 2021

  13. [21]

    BindCraft: one-shot design of functional protein binders.bioRxiv, pages 2024–09, 2024

    Martin Pacesa, Lennart Nickel, Christian Schellhaas, Joseph Schmidt, Ekaterina Pyatova, Lucas Kissling, Patrick Barendse, Jagrity Choudhury, Srajan Kapoor, Ana Alcaraz-Serna, et al. BindCraft: one-shot design of functional protein binders.bioRxiv, pages 2024–09, 2024

  14. [22]

    bioRxiv,pages2025–04,2025

    YehlinCho,MartinPacesa,ZhidianZhang,BrunoECorreia,andSergeyOvchinnikov.Boltzdesign1:Inverting all-atomstructurepredictionmodelforgeneralizedbiomolecularbinderdesign. bioRxiv,pages2025–04,2025

  15. [23]

    De novo design of protein structure and function with RFdiffusion.Nature, 620(7976):1089–1100, 2023

    JosephLWatson,DavidJuergens,NathanielRBennett,BrianLTrippe,JasonYim,HelenEEisenach,Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with RFdiffusion.Nature, 620(7976):1089–1100, 2023

  16. [24]

    De novo design of high-affinity protein binders with AlphaProteo.arXiv preprint arXiv:2409.08022, 2024

    ViniciusZambaldi,DavidLa,AlexanderEChu,HarshniraPatani,AmyEDanson,TristanOCKwan,Thomas Frerix, Rosalia G Schneider, David Saxton, Ashok Thillaisundaram, et al. De novo design of high-affinity protein binders with AlphaProteo.arXiv preprint arXiv:2409.08022, 2024

  17. [25]

    Zero-shot antibody design in a 24-well plate.bioRxiv, 2025

    Chai Discovery Team, Jacques Boitreaud, Jack Dent, Danny Geisz, Matthew McPartlon, Joshua Meier, Zhuoran Qiao, Alex Rogozhnikov, Nathan Rollins, Paul Wollenhaupt, and Kevin Wu. Zero-shot antibody design in a 24-well plate.bioRxiv, 2025

  18. [26]

    Accurate de novo design of high-affinity protein-binding macrocycles using deep learning.Nature Chemical Biology, pages 1–9, 2025

    StephenARettie,DavidJuergens,VictorAdebomi,YensiFloresBueso,QinqinZhao,AlexandriaNLeveille, Andi Liu, Asim K Bera, Joana A Wilms, Alina Üffing, et al. Accurate de novo design of high-affinity protein-binding macrocycles using deep learning.Nature Chemical Biology, pages 1–9, 2025

  19. [27]

    Robust deep learning–based protein sequence design using ProteinMPNN.Science, 378(6615):49–56, 2022

    JustasDauparas,IvanAnishchenko,NathanielBennett,HuaBai,RobertJRagotte,LukasFMilles,BasileIM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learning–based protein sequence design using ProteinMPNN.Science, 378(6615):49–56, 2022

  20. [28]

    Improving de novo protein binder design with deep learning

    Nathaniel R Bennett, Brian Coventry, Inna Goreshnik, Buwei Huang, Aza Allen, Dionne Vafeados, Ying Po Peng, Justas Dauparas, Minkyung Baek, Lance Stewart, et al. Improving de novo protein binder design with deep learning. Nature Communications, 14(1):2625, 2023

  21. [29]

    The protein data bank.Nucleic acids research, 28(1):235–242, 2000

    Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. The protein data bank.Nucleic acids research, 28(1):235–242, 2000

  22. [30]

    AlphaFold protein structure database Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design 18 in2024:Providingstructurecoverageforover214millionproteinsequences

    Mihaly Varadi, Damian Bertoni, Paulyna Magana, Urmila Paramval, Ivanna Pidruchna, Malarvizhi Radhakr- ishnan, Maxim Tsenkov, Sreenath Nair, Milot Mirdita, Jingi Yeo, et al. AlphaFold protein structure database Latent-X: An Atom-level Frontier Model for De Novo Protein Binder D...

  23. [31]

    Rational approaches to improving selectivity in drug design

    David J Huggins, Woody Sherman, and Bruce Tidor. Rational approaches to improving selectivity in drug design. Journal of medicinal chemistry, 55(4):1424–1444, 2012

  24. [32]

    A computationally designed inhibitor of an Epstein-Barr viral Bcl-2 protein induces apoptosis in infected cells.Cell, 157(7):1644–1656, 2014

    Erik Procko, Geoffrey Y Berguig, Betty W Shen, Yifan Song, Shani Frayo, Anthony J Convertine, Daciana Margineantu, Garrett Booth, Bruno E Correia, Yuanhua Cheng, et al. A computationally designed inhibitor of an Epstein-Barr viral Bcl-2 protein induces apoptosis in infected ce...

  25. [33]

    Targeting the MDM2-p53 interaction for cancer therapy.Clinical Cancer Research, 14(17):5318–5324, 2008

    Sanjeev Shangary and Shaomeng Wang. Targeting the MDM2-p53 interaction for cancer therapy.Clinical Cancer Research, 14(17):5318–5324, 2008

  26. [34]

    Targeting MCL-1 protein to treat cancer: opportunities and challenges.Frontiers in oncology, 13:1226289, 2023

    Shady I Tantawy, Natalia Timofeeva, Aloke Sarkar, and Varsha Gandhi. Targeting MCL-1 protein to treat cancer: opportunities and challenges.Frontiers in oncology, 13:1226289, 2023

  27. [35]

    De novo design of protein interactions with learned surface fingerprints.Nature, 617(7959):176–184, 2023

    Pablo Gainza, Sarah Wehrle, Alexandra Van Hall-Beauvais, Anthony Marchand, Andreas Scheck, Zander Harteveld, Stephen Buckley, Dongchun Ni, Shuguang Tan, Freyr Sverrisson, et al. De novo design of protein interactions with learned surface fingerprints.Nature, 617(7959):176–184, 2023

  28. [36]

    Preclinicalproofofprinciplefororally delivered Th17 antagonist miniproteins.Cell, 187(16):4305–4317, 2024

    StephanieBerger,FranziskaSeeger,Ta-YiYu,MerveAydin,HuilinYang,DanielRosenblum,LaureGuenin- Macé,CalebGlassman,LaurenArguinchona,CatherineSniezek,etal. Preclinicalproofofprinciplefororally delivered Th17 antagonist miniproteins.Cell, 187(16):4305–4317, 2024

  29. [37]

    Antagonism of nerve growth factor-TrkA signaling and the relief of pain.Anesthesiology, 115(1):189, 2011

    Patrick W Mantyh, Martin Koltzenburg, Lorne M Mendell, Leslie Tive, and David L Shelton. Antagonism of nerve growth factor-TrkA signaling and the relief of pain.Anesthesiology, 115(1):189, 2011

  30. [38]

    Cyclic peptide structure prediction and design using AlphaFold2.Nature Communications, 16(1):1–15, 2025

    StephenARettie,KatelynVCampbell,AsimKBera,AlexKang,SimonKozlov,YensiFloresBueso,Joshmyn De La Cruz, Maggie Ahlrichs, Suna Cheng, Stacey R Gerben, et al. Cyclic peptide structure prediction and design using AlphaFold2.Nature Communications, 16(1):1–15, 2025

  31. [39]

    Rational design of a potent macrocyclic peptide inhibitor targeting the PD-1/PD-L1 protein–protein interaction.RSC advances, 11(38): 23270–23279, 2021

    Qi Miao, Wanheng Zhang, Kuojun Zhang, He Li, Jidong Zhu, and Sheng Jiang. Rational design of a potent macrocyclic peptide inhibitor targeting the PD-1/PD-L1 protein–protein interaction.RSC advances, 11(38): 23270–23279, 2021

  32. [40]

    Discovery of cyclic peptide inhibitors targeting pd-l1 for cancer immunotherapy.Journal of medicinal chemistry, 65(18):12002–12013, 2022

    JohnFetse,ZhenZhao,HaoLiu,Umar-FaroukMamani,BahaaMustafa,PratikAdhikary,MohammedIbrahim, Yanli Liu, Pratikkumar Patel, Maryam Nakhjiri, et al. Discovery of cyclic peptide inhibitors targeting pd-l1 for cancer immunotherapy.Journal of medicinal chemistry, 65(18):12002–12013, 2022

  33. [41]

    MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets.Nature biotechnology, 35(11):1026–1028, 2017

    Martin Steinegger and Johannes Söding. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets.Nature biotechnology, 35(11):1026–1028, 2017

  34. [42]

    Fast and accurate protein structure search with foldseek

    Michel Van Kempen, Stephanie S Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron LM Gilchrist, Johannes Söding, and Martin Steinegger. Fast and accurate protein structure search with foldseek. Nature biotechnology, 42(2):243–246, 2024

  35. [43]

    Language models generalize beyond natural proteins

    Robert Verkuil, Ori Kabeli, Yilun Du, Basile IM Wicky, Lukas F Milles, Justas Dauparas, David Baker, Sergey Ovchinnikov, Tom Sercu, and Alexander Rives. Language models generalize beyond natural proteins. BioRxiv, pages 2022–12, 2022

  36. [44]

    Mammaliancelldisplayforantibodyengineering

    MitchellHoandIraPastan. Mammaliancelldisplayforantibodyengineering. Methods in Molecular Biology, 525:337–352, xiv, 2009

  37. [45]

    Beyond affinity: Selection of antibody variants with optimal biophysical properties and reduced immunogenicity from mammalian display libraries

    MichaelRDyson,EdwardMasters,DeividasPazeraitis,RajikaLPerera,JohannaLSyrjanen,SachinSurade, Nels Thorsteinson, Kothai Parthiban, Philip C Jones, Maheen Sattar, et al. Beyond affinity: Selection of antibody variants with optimal biophysical properties and reduced immunogenicity...

  38. [46]

    Macromolecular crystallographic information file (mmCIF)

    PhilipEBourne,HelenMBerman,BrianMcMahon,KeithDWatenpaugh,JohnDWestbrook,andPaulaMD Fitzgerald. Macromolecular crystallographic information file (mmCIF). InMethods in enzymology, volume 277, pages 571–590. Elsevier, 1997

  39. [47]

    lDDT:alocalsuperposition-free score for comparing protein structures and models using distance difference tests.Bioinformatics, 29(21): 2722–2728, 2013

    ValerioMariani,MarcoBiasini,AlessandroBarbato,andTorstenSchwede. lDDT:alocalsuperposition-free score for comparing protein structures and models using distance difference tests.Bioinformatics, 29(21): 2722–2728, 2013

  40. [48]

    Accurate structure prediction of biomolecular interactions with AlphaFold 3.Nature, pages 1–3, 2024

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3.Nature, pages 1–3, 2024

  41. [49]

    Chai-1: Decoding the molecular interactions of life.BioRxiv, 2024

    Chai Discovery team, Jacques Boitreaud, Jack Dent, Matthew McPartlon, Joshua Meier, Vinicius Reis, Alex Rogozhonikov, and Kevin Wu. Chai-1: Decoding the molecular interactions of life.BioRxiv, 2024

  42. [50]

    Boltz-2: Towards accurate and efficient binding Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design 19 affinity prediction

    Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, et al. Boltz-2: Towards accurate and efficient binding Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design...

  43. [51]

    How significant is a protein structure similarity with TM-score= 0.5? Bioinformatics, 26(7):889–895, 2010

    Jinrui Xu and Yang Zhang. How significant is a protein structure similarity with TM-score= 0.5? Bioinformatics, 26(7):889–895, 2010

  44. [52]

    Uniref clusters: A comprehensive and scalable alternative for improving sequence similarity searches

    Baris E Suzek, Yuqi Wang, Hongzhan Huang, Peter B McGarvey, Cathy H Wu, and UniProt Consor- tium. Uniref clusters: A comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, 2015

  45. [53]

    Biopython: freely available python tools for computational molecular biology and bioinformatics.Bioinformatics, 25(11):1422–1423, 2009

    PeterJACock,TiagoAntao,JeffreyTChang,BradAChapman,CymonJCox,AndrewDalke,IddoFriedberg, Thomas Hamelryck, Frank Kauff, Bartek Wilczynski, et al. Biopython: freely available python tools for computational molecular biology and bioinformatics.Bioinformatics, 25(11):1422–1423, 2009

  46. [54]

    Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.Biopolymers: Original Research on Biomolecules , 22(12): 2577–2637, 1983

    Wolfgang Kabsch and Christian Sander. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.Biopolymers: Original Research on Biomolecules , 22(12): 2577–2637, 1983

  47. [55]

    Codon optimization tool

    Twist Bioscience. Codon optimization tool. https://www.twistbioscience.com/resources/digital-tools/ codon-optimization-tool, 2025. Accessed: 2025-07-16. Supplementary information A. Model details for users Latent-X is available athttps://platform.latentlabs.com. In the followi...

  48. [56]

    The target can consist of multiple protein chains or crops of protein chains

    Target structure:The mmCIF file that describes the structure and sequence of the target protein(s) for which binders are designed. The target can consist of multiple protein chains or crops of protein chains

  49. [57]

    At least one hotspot needs to be provided

    Hotspot residues:The sequence location of the subset of target residues that constitute the target’s binding hotspots. At least one hotspot needs to be provided. In practice a small number of surface accessible and spatially close hotspots suffices and hotspots can be effectiv...

  50. [58]

    Binderlength: Thesequencelengthofthebindertobegenerated,measuredinnumberofaminoacidresidues

  51. [59]

    Targetcropping: Latent-Xcanbeconditionedoncroppedtargets,forexampleinordertofitthecontextwindow. The user should be aware of the following details on input representations: • Context length:The context length for inputs and outputs is 512 residues, counting jointly residues in...

  52. [60]

    hingeeffect

    and Boltz-2 [50]. These models reproduce the AlphaFold 3 architecture and approach AlphaFold 3 protein structure prediction performance. We used Chai-1 based filtering in our experiments, but show in App. D.3 that Chai-1 and Boltz-2 have qualitatively similar filter performanc...

  53. [61]

    Gen- Script

    and SC2RBD [7], and the best RFdiffusion designs for IL-7R𝛼, PD-L1, and TrkA [23] as positive controls in our mDisplay and BLI measurements seen in Tab. S5. For BLI, we also included the best binders for each target from AlphaProteo that represent the previous best-in-class de...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.