{"id":"f3e6a417-2afa-4227-80cf-ee8aecfe20b9","arxiv_id":"2507.19375","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A proprietary all-atom generative model, Latent-X, designs protein binders that achieved high experimental hit rates (91-100% for macrocycles, 10-64% for mini-binders) across seven targets, with affinities down to picomolar levels.","lead":"Latent-X is an AI model that designs protein binders from a target structure, and its designs showed high experimental success rates: macrocyclic peptides bound their targets in 91-100% of tests and mini-binders reached low nanomolar and picomolar affinities. The paper compares the model head-to-head with AlphaProteo, RFdiffusion and RFpeptides, reports an in silico study on 200 held-out targets, and claims roughly 10x faster generation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Macrocycle 'as few as 30 designs' and >90% hit-rate claims rest on a post-synthesis, post-filter subset of 11–17 tested designs, not on 30 submitted designs.","rationale":"The paper's strongest claim combines three quantitative assertions: macrocycle hit rates >90%, successful design with 30–100 tested designs, and superiority over prior methods. The first two are load-bearing for the 'push-button' framing. My review of Sec. 3.1 and Table 2 shows that the macrocycle hit rate is calculated on only 11–17 SPR-tested designs per target, selected from the 30 that survived synthesis and from the top-30-of-700 after in silico filtering. A hit rate of 90.9% is 10/11. With such denominators, the difference between a 90% and a 60% true hit rate is within sampling error (exact binomial 95% CI for 10/11 is roughly 59–100%), so the 'exceeding 90% on all targets' claim is fragile. The same selection pipeline, not raw generation, is what is being evaluated. This is not an accusation of misconduct; the paper discloses the filtering and the post-hoc relaxation in App. B.1, but the abstract and headline do not carry the denominator caveat. The reader's weakest_assumption about filter reliability is related, but my focus is the denominator and selection, which is more concrete and directly testable. The mini-binder results are stronger because 100 designs per target were tested, so the central claim is not fully undermined. The recommended verdict remains CONDITIONAL: the macrocycle wording and denominator need correction, and independent verification of the selection pipeline would be required before the headline claim can be accepted as stated.","tokens_in":33732,"tokens_out":7373,"duration_ms":74085,"concrete_test":"Recompute the macrocycle hit rate with an intention-to-test denominator: take the 30 designs per target that were submitted for synthesis, count synthesis failures as non-hits or as non-evaluable in two sensitivity analyses, and report the hit rate over the full 30. If the >90% figure drops materially when synthesis failures are included, the abstract wording must be revised. As a second confirmatory check, randomly sample 30 of the 700 generated macrocycles for one target, bypass the in silico filter, and test them in the same SPR assay; if the unfiltered hit rate is substantially below the reported 94–100%, the headline result should be attributed to the filter-plus-selection pipeline rather than to raw generation quality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline experimental claim that macrocycles are validated with 'as few as 30–100 designs' and hit >90% is not supported by the reported denominator. Section 3.1 states that 30 macrocycles per target were submitted for synthesis, but only 57–87% cyclized at >90% purity, and only 11–17 designs per target were then selected for SPR binding assessment. The hit rates in Table 2 (90.9% for MDM2, 100% for MCL-1, 94.1% for PD-L1) are computed on that selected subset, not on the 30 submitted designs, and not on the 700 generated per target. A 91% hit rate is 10/11 binders, not 27/30. In addition, the 11–17 tested designs were chosen by an in silico macrocycle filter whose ptm_binder component was dropped post hoc because known macrocycle binders failed it (App. B.1), and by a top-30-of-700 ranking. Thus the >90% hit rate is a property of the full filter-plus-selection pipeline applied to a small, non-random subset, and the abstract's 'as few as 30–100 designs' wording overstates how many designs were actually tested. The practical claim that Latent-X reliably produces functional binders with 30–100 lab tests is therefore not established for macrocycles.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Latent-X is an all-atom generative model that, given a target structure and hotspot residues, co-generates the structure and sequence of a protein binder together with the bound target conformation. The authors validate the full pipeline (generation plus in silico filtering) in wet-lab experiments across two modalities: macrocyclic peptides (12–18 residues) against MDM2, MCL-1, and PD-L1, and mini-binders (80–120 residues) against BHRF1, IL-7Ra, PD-L1, SC2RBD, and TrkA. Headline claims are experimental hit rates of 91–100% for macrocycles (computed on 11–17 SPR-tested designs per target) and 10–64% for mini-binders (100 HT-BLI-tested per target, with reported KD values down to <0.01 nM for several designs), specificity in an orthogonal mammalian-display assay, an order-of-magnitude inference speed advantage over RFdiffusion, and higher expected in silico hit rates than RFpeptides/RFdiffusion on 200 held-out PDB targets (8.26% vs 1.72% for macrocycles; 5.11% vs 3.02% for mini-binders). The paper also reports replicated affinity measurements of previously published AlphaProteo, RFdiffusion, and RFpeptides binders under the same assay conditions.","tokens_in":33945,"tokens_out":20131,"duration_ms":168459,"significance":"If the mini-binder results are taken as reported, they represent a genuinely strong empirical contribution: 100 designs per target were tested with a stated response threshold (0.03 RU), hit rates range from 10% to 64%, affinities were confirmed by 5-point BLI on a subset, and an orthogonal mDisplay assay correlates with HT-BLI (Pearson r = 0.68–0.79). The prospective in silico benchmark on 200 structures deposited after the training cutoff of both the design and structure-prediction models, using openly available Chai-1/Boltz-2 metrics and repeated for both filters, is a serious attempt at estimating generalization, and the paper provides the data needed to check the reported funnel. Credit is also due for the explicit disclosure of the post-hoc macrocycle filter relaxation (App. B.1), the mini-binder filter bias (Sec. 5.3), and synthesis attrition (Sec. 3.1), and for reporting the weaker macrocycle KD values in Fig. S8 rather than only the best ones.","major_comments":[{"comment":"The headline macrocycle claim is not supported by the reported data. Section 3.1 states that 30 designs per target were submitted for synthesis, that only 57% (MDM2), 77% (MCL-1), and 87% (PD-L1) cyclized at >90% purity, and that 11–17 designs per target were selected for SPR. The hit rates in Table 2 (90.9% for MDM2, 100% for MCL-1, 94.1% for PD-L1) are therefore 10/11, 11/11, and 16/17, computed on a small, non-random subset of the submitted designs. The selection criteria for this subset ('chosen to provide a representative sample') are not specified, so the abstract's 'testing as few as 30–100 designs per target' and 'hit rates exceeding 90%' are misleading; even under the optimistic assumption that every untested cyclized design binds, the per-submitted-design ceilings are 33–53%. Second, the hit definition ('measurable binding', Table 2 footnote) counts SPR designs whose fitted KD is in the millimolar range, several orders of magnitude above the highest analyte concentration used (25–500 µM, Fig. S12), for example LL_CYC_PD-L1_22 at 10.4 mM and LL_CYC_MCL-1_16 at 1.34 mM in Fig. S8; these are extrapolated fits, not established binding. Applying a KD < 100 µM threshold to the values in Fig. S8 reduces the hit rates to roughly 5/11 for MDM2, 6–7/11 for MCL-1, and 3–4/17 for PD-L1. Please report the full funnel (generated, filter-passed, submitted, cyclized, tested, bound at an explicit threshold) and revise the abstract and Sec. 5.1 accordingly.","section":"Sec. 3.1, Table 2, Fig. S8; Abstract"},{"comment":"The claim of 'direct comparisons ... under identical conditions ... higher hit rates' (abstract) is contradicted by the paper's own table. Table 2's footnote states that AlphaProteo and RFdiffusion/RFpeptides hit rates are 'taken from the original publications,' which used different assays and thresholds (e.g., AlphaProteo's yeast-display pipeline versus the present HT-BLI); only the affinity measurements in Table 1, on replicated binder sequences, were made under identical conditions. Furthermore, the paper's own Table 2 shows AlphaProteo's published hit rate of 88% on BHRF1 exceeding Latent-X's 64%, and RFdiffusion's published 33.7% on IL-7Ra exceeding Latent-X's 26.0%. The blanket statements in the abstract and in Sec. 5.2 ('consistent advantages across all tested targets') therefore overstate the evidence. Please either restrict the hit-rate claims to the targets and assays where they hold, or report measured hit rates for the comparator workflows under the same assay and threshold.","section":"Abstract, Table 2, Sec. 5.2"},{"comment":"The macrocycle in silico filter was modified after observing failures. App. B.1 reports that the ptm_binder criterion was dropped because three known experimentally validated macrocycle complexes (9cdz, 7oun, 1sfi) failed it, and that 'this choice of filter for macrocycles was post-hoc justified by the high resulting experimental success.' Because the 11–17 SPR-tested macrocycles per target were selected from the top-30 of 700 by this relaxed filter, the >90% hit rates are conditional on a filter whose key modification was motivated by the outcome the filter is then used to validate. The paper should (i) report how many of the 11–17 tested macrocycles would have passed the unmodified filter, (ii) present the relaxation as a limitation at the level of the abstract rather than only in an appendix, and (iii) attribute the reported rate to the full model-plus-filter selection pipeline rather than to the generative model alone.","section":"App. B.1, Sec. 3.1"}],"minor_comments":[{"comment":"The caption states '25 de novo designed bound structures across 3 different targets,' but the figure lists 37 structures (10 MDM2, 11 MCL-1, 16 PD-L1); the caption should be corrected.","section":"Fig. S8"},{"comment":"KD values reported as '<0.01 nM' (e.g., LL_MINI_IL-7Ra_39, LL_MINI_SC2RBD_8) are below the reliable quantification range of the 5-point BLI assay and should be reported as upper bounds (KD < 10 pM) to avoid overstating precision.","section":"Sec. 3.2, Tab. S4"},{"comment":"The sentence beginning 'In direct comparisons with the state-of-the-art models ... under identical conditions demonstrates' is grammatically incomplete and should be rewritten.","section":"Abstract"},{"comment":"The experimental results are from Latent-X v1, whereas the in silico expected hit-rate benchmark (Fig. 6) uses v1.1; this version mismatch should be stated wherever Fig. 6 results are cited, since the reader cannot judge whether the two sets of numbers describe the same model.","section":"Sec. 2.1 vs Sec. 4"},{"comment":"The target name is rendered as 'IL-7Rα' in some places and 'IL-7Ra' in others; the notation should be unified.","section":"Sec. 3.1, Fig. 5, Tab. S3"},{"comment":"The number of Tier-2 designs per target advanced from HT-BLI to 5-point BLI, and the precise selection rule (thresholds on Ka and RU), are not reported; without these, the KD distribution in Fig. S9 cannot be related to the HT-BLI hit rates.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is authored exclusively by employees of, or advisors to, Latent Labs, and the declared competing interest is prominent; the text also includes promotional phrasing ('push-button biologics discovery', 'democratizes advanced protein design') that should be tempered in a journal version. The model itself is proprietary: no architecture, training objective, data composition, or hyperparameter details are given, so the scientific content is almost entirely empirical. That may be acceptable for this venue only if the empirical record is reported without spin; currently the abstract and Sec. 5 overstate the macrocycle results and the 'identical conditions' comparison. The editor may also wish to consider whether a fully unreviewable model with a live commercial platform is within scope, or whether the extensive wet-lab validation is sufficient. In my view the mini-binder evidence and the prospective in silico benchmark are strong enough to warrant a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading. The paper reports genuinely new wet-lab binder designs with sequences, structures, and affinities, plus a thoughtful 200-target in silico benchmark. But the headline numbers need careful reading: the >90% macrocycle hit rates are computed on 11–17 selected designs per target, not on the 30 submitted or the 700 generated, so \"as few as 30–100 designs\" is not quite what the data show.\n\nWhat's new: Latent-X generates macrocycles and mini-binders zero-shot, including beta-sheet folds that RFdiffusion doesn't produce; the all-atom co-generation of binder and target is a real design choice. The experimental work is substantial: 87 mini-binder and 25 macrocycle Kd values, orthogonal validation with BLI, SPR, and mDisplay, and sequences provided in Tab S4 for external follow-up. The in silico benchmark on 200 PDB entries released after the training cutoff, with filters tuned on the independent Cao et al. set, is a sensible protocol and a real comparison against RFdiffusion/RFpeptides.\n\nSoft spots, in order of importance. First, the macrocycle hit-rate denominator. Only 11–17 of 30 submitted designs per target made it through synthesis, cyclization, and purity selection; the 91–100% hit rates are on that subset. That's 10/11 for MDM2, 17/17 for MCL-1, and 16/17 for PD-L1 — decent evidence, but not \"testing 30 designs.\" Second, the filters are part of the pipeline. The macrocycle filter dropped ptm_binder after known binders failed it (App B.1), and the filter thresholds were tuned on helical 45–65 aa mini-binders. The authors disclose this bias, and it means the high hit rates reflect the full filter-plus-selection workflow, not the model alone. Third, the model is proprietary: no architecture, weights, or code, so the generative claims cannot be independently reproduced. That doesn't invalidate the wet-lab results, but it limits the paper to a platform claim. Fourth, the comparison numbers are slippery: published vs replicated Kd values for AlphaProteo and RFdiffusion differ a lot, and the abstract says \"under identical conditions\" when only the replicated column is truly that.\n\nNone of this kills the central empirical result. Given the sequences and affinities, the evidence that Latent-X generated working binders is solid. The paper's own limitations section is honest. But the abstract and discussion push the \"30–100 designs\" and \"exceeding 90%\" wording further than the data support.\n\nThis deserves serious peer review with major revision: fix the denominators, qualify the filter bias, provide per-design errors, and ideally release the model or a detailed architecture. As a reader, I'd cite the empirical sequences and the benchmark protocol.","headline":"Real wet-lab results and a credible 200-target benchmark, but the headline macrocycle hit rates rest on small post-filter subsets and the model is proprietary, so the 'as few as 30–100 designs' claim overstates what is actually demonstrated.","tokens_in":34636,"tokens_out":2289,"would_cite":true,"duration_ms":22266,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Latent-X jointly designs protein binder sequence and all-atom structure, reporting 91-100% macrocycle hit rates and picomolar affinities in wet-lab tests.","keywords":["de novo protein binder design","all-atom generative model","macrocyclic peptides","mini-binders","epitope conditioning","structure prediction filtering","wet-lab validation","binding affinity"],"falsifier":"Test the claim by taking a panel of, say, 20 target proteins unrelated to the filter-tuning set, running the complete Latent-X workflow, and synthesising a random sample of both pass and fail designs from the in silico filter; the paper's claim predicts a large gap in experimental hit rates between pass and fail sets, while the filter-bias objection predicts comparable rates, especially on non-helical folds. A simpler check: for one macrocycle target, synthesise the roughly 600 designs that did not pass the filter and measure what fraction still bind, since a >90% generation-level hit rate predicts many should.","tokens_in":33448,"feed_emoji":"🧬","tokens_out":8212,"duration_ms":71527,"temperature":0.7,"pith_summary":"This paper introduces Latent-X, a generative model that takes a target protein and a few hotspot residues and produces both the amino-acid sequence and the all-atom structure of a protein binder, while also co-generating the target side chains at the interface. The central claim is that binder design can be reduced to testing a few dozen to a hundred designs per target rather than screening millions of molecules. In wet-lab experiments across seven targets the paper reports macrocycle hit rates from 91% to 100%, mini-binder hit rates from 10% to 64%, and affinities reaching below 0.01 nanomolar, with the best numbers beating previously reported binders from the strongest existing methods under identical assay conditions. The paper further claims that the model runs an order of magnitude faster than multi-step pipelines and produces structurally diverse binders, including beta-sheet folds.","feed_headline":"Protein binders designed by AI hit 90-100% in wet-lab tests","feed_subtitle":"With only 30-100 designs per target, the model reaches picomolar affinities and beats prior design pipelines.","key_machinery":"The load-bearing mechanism is joint co-generation of sequence and all-atom structure for binder and target in a single generative pass, in contrast to conventional pipelines that first sample a backbone and then assign a sequence with a separate model; this lets the model optimise interface side-chain rotamers and hydrogen-bonding networks directly. Around the generator sits an in silico filter built on structure-prediction models, using thresholds on the minimum interface predicted aligned error (min_ipae), the binder's predicted TM-score (ptm_binder) and the RMSD between generated and predicted complex (complex_rmsd), which selects the handful of designs that go to the lab.","core_discovery":"Latent-X is an all-atom generative model that simultaneously predicts the sequence and full atomic structure of both the binder and the target protein in one pass. Prompted with a target structure and hotspot residues, it directly constructs non-covalent interactions, including extensive interface hydrogen-bond networks, instead of leaving interface chemistry to post-hoc refinement. In the authors' experimental validation, macrocycles against MDM2, MCL-1 and PD-L1 achieved hit rates of 90.9%, 100% and 94.1% respectively with best affinities of 5.35 micromolar, 18.4 micromolar and 71.7 micromolar, and mini-binders against BHRF1, TrkA, PD-L1, IL-7R-alpha and SC2RBD achieved hit rates of 64%, 10%, 49%, 26% and 52% with best affinities ranging from 32.4 nanomolar on BHRF1 to below 0.01 nanomolar on IL-7R-alpha and SC2RBD. These results are presented as outperforming the best published binders from RFdiffusion, RFpeptides and AlphaProteo when replicated in identical assays, with an order-of-magnitude faster generation and an in silico hit-rate advantage on 200 held-out PDB targets.","pith_inferences":["A consequence the authors leave implicit: the reported macrocycle hit rates are for the full pipeline of generator plus relaxed macrocycle filter, so the generator's standalone success rate on unfiltered designs is probably lower; the paper's own in silico success rates of 56-67% before synthesis filters hint at the magnitude.","The disclosed filter bias toward 45-65 residue helical mini-binders suggests the 10% hit rate on TrkA and 26% on IL-7R-alpha may understate what the model itself achieves on non-helical folds; testing unfiltered or filter-agnostic sets would quantify this.","A testable extension: because conditioning is via hotspot residues, the same pipeline could be asked to design binders with prescribed multi-target specificity or to hit two epitopes on one target, which the paper only gestures at in future work.","The in silico benchmark on 200 held-out PDB targets predicts that some fraction of arbitrary targets, roughly 5.5% for mini-binders by the paper's own count, produce no passing designs; a prospective user should expect occasional total failures and plan around them."],"forward_implications":["If the wet-lab numbers hold, a researcher can go from a target structure to candidates worth testing with 30-100 designs per target instead of millions of screened molecules.","Head-to-head superiority under identical conditions would make Latent-X the default generative method for macrocyclic peptide and mini-binder design, displacing backbone-diffusion-plus-resequencing workflows.","An order-of-magnitude inference speed-up makes design interactive and allows large in silico sweeps over targets and epitopes as a routine step.","Structural diversity, including beta-sheet folds, suggests the accessible design space extends beyond the helical bundles that previous methods mainly produced.","Zero-shot generalisation to both cyclic and linear topologies without task-specific fine-tuning implies one model can serve multiple therapeutic modalities."],"supporting_citations":[{"why":"Supplies the experimentally characterised binder-target dataset on which the in silico filters were tuned, and the prior design-from-target-structure benchmark.","marker":"[9]"},{"why":"Defines the RFdiffusion mini-binder workflow and experimental benchmarks that Latent-X is compared against.","marker":"[23]"},{"why":"Defines the AlphaProteo benchmark, the filter-tuning procedure and thresholds that Latent-X adopts and claims to beat.","marker":"[24]"},{"why":"Provides the RFpeptides macrocycle benchmark and the reported hit rates that Latent-X claims to exceed.","marker":"[26]"},{"why":"ProteinMPNN is the sequence-design step in the baseline workflows, used here to define the comparison pipeline.","marker":"[27]"},{"why":"Documents the filter-tuning methodology for improving de novo binder design that Latent-X's filters follow.","marker":"[28]"},{"why":"Provides macrocycle design benchmarks, positive controls and comparison hit rates for MDM2 and MCL-1.","marker":"[38]"},{"why":"Chai-1, the structure-prediction model behind the primary in silico filter used to select designs.","marker":"[49]"},{"why":"Boltz-2, the alternative structure-prediction model used to show filter results do not depend on the choice of prediction model.","marker":"[50]"},{"why":"AlphaFold 2, whose structure-quality metrics and relax protocol are used to validate the stereochemical quality of Latent-X designs.","marker":"[10]"}],"fun_headline_variants":["Latent-X designs protein binders with up to 100% hit rates in wet labs","AI model Latent-X: 90%+ binder hits, picomolar affinities","One-pass all-atom design yields potent binders, beating AlphaProteo","Latent-X: from epitope to picomolar binder in a single pass","AI binder design: 90-100% hit rates with just 30-100 candidates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the in silico filters built on structure-prediction confidence metrics pick out designs that will actually bind in the lab; the paper itself concedes these filters were tuned on helical 45-65 residue mini-binders and that the macrocycle filter had to be relaxed after known binders failed it, so if the filters select for easy-to-predict rather than genuinely potent proteins, the reported hit rates overstate the model's stand-alone generation quality.","fun_headline_variants_meta":{"raw":{"variants":["Latent-X designs protein binders with up to 100% hit rates in wet labs","AI model Latent-X: 90%+ binder hits, picomolar affinities","One-pass all-atom design yields potent binders, beating AlphaProteo","Latent-X: from epitope to picomolar binder in a single pass","AI binder design: 90-100% hit rates with just 30-100 candidates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000415,"raw_usage":{"total_tokens":2235,"prompt_tokens":1128,"completion_tokens":1107,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":744,"completion_tokens_details":{"reasoning_tokens":995}},"tokens_in":744,"tokens_out":1107,"duration_ms":8586,"temperature":1.0,"reasoning_tokens":995,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:54:11.721303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Test the claim by taking a panel of, say, 20 target proteins unrelated to the filter-tuning set, running the complete Latent-X workflow, and synthesising a random sample of both pass and fail designs from the in silico filter; the paper's claim predicts a large gap in experimental hit rates between pass and fail sets, while the filter-bias objection predicts comparable rates, especially on non-helical folds. A simpler check: for one macrocycle target, synthesise the roughly 600 designs that did not pass the filter and measure what fraction still bind, since a >90% generation-level hit rate predicts many should.","supporting_citations":[{"cited_title":"Design of protein-binding proteins from the target structure alone.Nature, 605(7910):551–560, 2022","cited_arxiv_id":null,"evidence_quote":"Supplies the experimentally characterised binder-target dataset on which the in silico filters were tuned, and the prior design-from-target-structure benchmark."},{"cited_title":"De novo design of protein structure and function with RFdiffusion.Nature, 620(7976):1089–1100, 2023","cited_arxiv_id":null,"evidence_quote":"Defines the RFdiffusion mini-binder workflow and experimental benchmarks that Latent-X is compared against."},{"cited_title":"Accurate de novo design of high-affinity protein-binding macrocycles using deep learning.Nature Chemical Biology, pages 1–9, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the RFpeptides macrocycle benchmark and the reported hit rates that Latent-X claims to exceed."},{"cited_title":"Robust deep learning–based protein sequence design using ProteinMPNN.Science, 378(6615):49–56, 2022","cited_arxiv_id":null,"evidence_quote":"ProteinMPNN is the sequence-design step in the baseline workflows, used here to define the comparison pipeline."},{"cited_title":"Cyclic peptide structure prediction and design using AlphaFold2.Nature Communications, 16(1):1–15, 2025","cited_arxiv_id":null,"evidence_quote":"Provides macrocycle design benchmarks, positive controls and comparison hit rates for MDM2 and MCL-1."},{"cited_title":"Chai-1: Decoding the molecular interactions of life.BioRxiv, 2024","cited_arxiv_id":null,"evidence_quote":"Chai-1, the structure-prediction model behind the primary in silico filter used to select designs."},{"cited_title":"Boltz-2: Towards accurate and efficient binding Latent-X: An Atom-level Frontier Model for De Novo Protein Binder Design 19 affinity prediction","cited_arxiv_id":null,"evidence_quote":"Boltz-2, the alternative structure-prediction model used to show filter results do not depend on the choice of prediction model."},{"cited_title":"Highly accurate protein structure prediction with AlphaFold.nature, 596(7873):583–589, 2021","cited_arxiv_id":null,"evidence_quote":"AlphaFold 2, whose structure-quality metrics and relax protocol are used to validate the stereochemical quality of Latent-X designs."}],"review_version":2}