{"id":"94290fa9-5325-40f3-902e-f9b795b2c088","arxiv_id":"2507.09054","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"Ibex predicts antibody, nanobody, and TCR structures with an explicit apo/holo conformation token, achieving competitive accuracy with much lower compute than baseline models.","lead":"Ibex is a new computer model that predicts the 3D shapes of antibodies, nanobodies, and T-cell receptors from their amino acid sequence. It can predict both the free (unbound) and bound shapes of the same protein, which matters for designing therapeutic antibodies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central apo/holo claim lacks a held-out evaluation: Section 3.1 states that most paired apo/holo structures were in training, so Figures 2A–C largely measure memorization; generalization to novel sequences is unverified.","rationale":"The reader's verdict is CONDITIONAL, and my concern supports that verdict rather than moving it. The weakest point is not primarily label noise in §5.2, although that matters. The decisive gap is that the paper's key evidence for the conformation token—Figures 2A–C—is produced on the 562 matched apo/holo pairs that the authors state were 'mostly included in the training.' A model can score well on such pairs by memorizing the two structures associated with each sequence, especially because matched pairs are explicitly upsampled during training (§5.3: 'paired apo/holo are upsampled by a factor 2'). No held-out cluster split is reported for these pairs, and the private OOD benchmark (Section 3.3) is not used to test whether the apo/holo token produces the correct state-specific structure. The label-noise concern is real but secondary: even with perfect labels, the current results would not establish that the model generalizes the apo/holo distinction. The absence of a conformation-token ablation further weakens the mechanistic claim, though the residual connection (§5.1) makes the token's propagation plausible. These issues are addressable, so CONDITIONAL remains appropriate: the authors should provide a held-out apo/holo evaluation and, ideally, a private-dataset state-matched comparison.","tokens_in":15480,"tokens_out":5834,"duration_ms":69137,"concrete_test":"Hold out all clusters containing the 562 matched apo/holo pairs from training (using the same MMseqs2 95% CDR-sequence clustering as §5.2), then recompute Figure 2A/B on the held-out subset; report per-state CDR H3 RMSD and the apo–holo RMSD difference. If the model's accuracy on held-out pairs is substantially worse than on the full set, or the apo and holo predictions collapse toward a single structure, the generalization claim fails. As a complementary check, run Ibex with the apo and holo tokens on the 286 private structures (177 holo, 109 apo) and test whether the token-matched prediction is closer to the experimental structure than the mismatched token.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Ibex explicitly distinguishes bound and unbound conformations and predicts both accurately from one sequence—is not supported by a held-out evaluation. Section 3.1 states: 'It is important to note here that most of the known paired apo/holo structures were included in the training.' The 562-pair analysis in Figures 2A–C therefore largely measures memorization, not generalization. No split is reported isolating held-out pairs, and no apo/holo evaluation is performed on the private 286-structure benchmark (Section 3.3), even though that set contains 177 holo and 109 apo structures (Appendix C). Consequently, the conformation token might simply implement a learned lookup for training pairs; its ability to produce distinct, accurate states for novel sequences is unverified. The label-noise issue (§5.2) is secondary: even with perfect labels, the current evidence cannot distinguish a generalizable conformational model from one that has memorized the paired training data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Ibex is a deep learning model for structure prediction of antibody, nanobody, and TCR variable domains. It builds on AlphaFold2/ABodyBuilder3 with a conformation token that is set to apo or holo during training and at inference. The model is trained in three stages on SAbDab/STCRDab structures, PDB immunoglobulin-like domains, and ESMFold/Boltz-1 predicted structures from OAS sequences. The paper validates the model on (i) 562 paired apo/holo antibody structures, (ii) the ImmuneBuilder test set, and (iii) a private set of 286 novel antibody structures, and reports a substantial speed advantage over diffusion-based general predictors.","tokens_in":15764,"tokens_out":4063,"duration_ms":46489,"significance":"The strongest contribution is the combination of released inference code and weights with a private, out-of-distribution benchmark of 286 novel antibody structures and an ablation study showing the value of the auxiliary training data. If the conformation-aware claim survives a held-out evaluation, the model would be a useful tool for antibody design. At present, however, the key novelty—predicting distinct and accurate apo and holo structures from a single sequence—is only demonstrated on structures that largely overlap the training set, so its generalization is not yet established.","major_comments":[{"comment":"The apo/holo validation is performed on 562 paired structures, and the paper explicitly states that 'most of the known paired apo/holo structures were included in the training.' Consequently, Figures 2A–C cannot distinguish memorization from generalization, and the central claim that Ibex predicts both conformations for novel sequences is not supported by this analysis. Provide a held-out evaluation, for example by splitting the 562 pairs according to the MMseqs2 cluster definition so that no pair sharing a cluster with a training structure is included, or by repeating the apo/holo analysis on the private dataset of Section 3.3, which per Appendix C contains 177 holo and 109 apo structures. Without such a split, the conformation token may simply implement a learned lookup for training pairs.","section":"Section 3.1"},{"comment":"The abstract's 'state-of-the-art accuracy' claim is not fully supported by Table 1: Chai-1 achieves a lower mean CDR H3 RMSD than Ibex on antibodies (2.65 Å vs 2.72 Å) and Boltz-1 achieves a lower mean CDR H3 RMSD on nanobodies (2.83 Å vs 3.12 Å). The claim should be qualified to the specific regions and molecule types where Ibex is actually best (e.g., TCR CDR β3 and CDR α3), and the comparison should include measures of variance or paired significance tests, since the reported differences are often small (e.g., 0.02–0.10 Å).","section":"Table 1 / Abstract"},{"comment":"The apo/holo label is assigned by the presence of an antigen chain in SAbDab/STCRDab metadata. This binary rule is likely to be noisy: a crystal structure can lack the antigen in the asymmetric unit, and crystallographic contacts can be non-physiological. Because the conformation token is trained directly on these labels, label noise may cause the model to learn spurious distinctions rather than genuine conformational states. The authors should validate a sample of the labels (e.g., manual inspection or comparison with structural metrics such as buried surface area or H3 loop conformation) and analyze the sensitivity of the apo/holo predictions to the labeling rule.","section":"Section 5.2"},{"comment":"The paper reports a 'reasonable correlation' between predicted and experimental conformational changes but gives no quantitative correlation coefficient, no mean/median loop RMSD values, and no confidence intervals. To support the claim that Ibex recapitulates conformational transitions, report Pearson and Spearman correlations between predicted and observed apo–holo CDR H3 RMSDs, the mean and median signed error, and the fraction of pairs for which the predicted direction of change agrees with experiment.","section":"Section 3.1, Figure 2A"}],"minor_comments":[{"comment":"The statement that Ibex shows 'comparable performance to Boltz-1' is based on a 0.02 Å difference in mean CDR H3 RMSD (2.28 vs 2.30 Å); report the distribution or a paired test to justify this comparison.","section":"Section 3.3"},{"comment":"The spelling 'SAbdab' appears in the first paragraph of Section 5.3 and should be 'SAbDab' for consistency with the rest of the manuscript.","section":"Section 5.3"},{"comment":"The caption of Figure 9 states that the comparison is against ABodyBuilder3, but the surrounding text says the comparison is to TCRBuilder2+; the caption should be corrected.","section":"Appendix B, Figure 9"},{"comment":"The y-axis of Figure 4 is labeled 'Relative improvement' but the plot does not specify whether higher values are better or how the single-seed baseline is defined; please add a precise definition in the caption.","section":"Figure 4"},{"comment":"The ensemble procedure 'returns the prediction closest to the mean' should specify the alignment and distance metric used (e.g., global backbone RMSD after superposition on framework residues).","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The central claim of conformation-aware prediction requires a held-out evaluation before the paper can be accepted; otherwise the abstract should be weakened to a training-set recapitulation result. The authors are affiliated with Genentech/Roche, and the private dataset is a key asset; the paper should state whether the private benchmark structures will be made available or deposited, since the OOD comparison cannot be independently reproduced otherwise. The manuscript is otherwise within the journal's scope and the release of code and weights is a positive factor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for two reasons: the conformation token is a genuinely new idea, and the evidence for its central claim is thinner than the abstract suggests. Ibex conditions on an apo/holo label and returns two distinct, accurate structures from a single variable-domain sequence. That is new. Prior antibody models give one static structure. The model is also fast, releases inference code and weights, and the authors are careful about leakage on the public ImmuneBuilder test set, even excluding Boltz-2 because it trained on most of it. The private 286-set benchmark with novel CDR H3 loops is a real out-of-distribution test for overall accuracy, and Ibex is competitive there, besting Boltz-1 on mean CDR H3 RMSD (2.28 vs 2.30).\n\nNow the soft spots, in order of real weight. First and most important: the flagship apo/holo capability is validated on 562 paired structures, and the paper states plainly that most of those pairs were in training. That means Figure 2 largely measures memorization, not generalization to new sequences. The authors admit this but do not supply a held-out apo/holo split. Nothing stops them: the private set contains 177 holo and 109 apo structures, but they never evaluate the conformation token there. So the central claim—that Ibex can generalize both conformational states to novel sequences—is unverified. This is a load-bearing gap, not a nitpick.\n\nSecond, the label-noise issue from the reader's report is real but secondary. The apo/holo rule relies on SAbDab/STCRDab metadata rather than physical contact, so some labels are noisy; however, even with perfect labels, the memorization problem would remain. Third, the abstract's \"state-of-the-art\" claim is too strong. Table 1 shows Chai-1 better on antibody CDR H3 (2.65 vs 2.72) and Boltz-1 better on nanobody CDR H3 (2.83 vs 3.12). Ibex is excellent on TCR loops and competitive across the board, but SOTA is an overstatement.\n\nWhat holds up: the architecture is a modest AlphaFold2 extension, but the capability is real and the engineering is solid. The authors are honest about limitations, including the binary state simplification and the scarcity of large conformational changes. The citation pattern looks fine.\n\nThis paper deserves a serious referee. The reviewer should ask for a held-out apo/holo evaluation (e.g., on the private set), a rejoinder to the training-overlap issue, and a toned-down abstract. The conformation token idea is worth publication even if the current evidence only proves memorization on known pairs; the field needs a model that can eventually do this. I'd bring it to reading group to discuss the token design and the evaluation gap. Cite it if you work on antibody structure prediction, since the code and weights are genuinely usable.","headline":"Ibex brings a genuinely new capability—predicting apo and holo antibody structures from one sequence via a conformation token—but the key generalization claim is validated on training pairs, and the abstract oversells state-of-the-art.","tokens_in":16301,"tokens_out":2142,"would_cite":true,"duration_ms":26117,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that explicitly labeling structures as apo or holo during training lets one model predict both the unbound and antigen-bound conformations of antibodies, nanobodies, and T-cell receptors from a single sequence.","keywords":["antibody structure prediction","conformation token","apo holo states","CDR H3 loop","T-cell receptor modeling","nanobody structure","out-of-distribution generalization","deep learning protein prediction"],"falsifier":"Take a set of antibody–antigen complexes crystallized both with and without the antigen under identical conditions, ideally newly deposited ones absent from training; if Ibex's apo and holo predictions do not separate into the experimentally observed states, or if shuffling the apo/holo labels during training leaves the two-state prediction gap unchanged, the binding-state conditioning is not capturing real conformational differences.","tokens_in":15333,"feed_emoji":"🧬","tokens_out":7402,"duration_ms":71753,"temperature":0.7,"pith_summary":"Ibex is a deep-learning model that predicts the three-dimensional structure of the variable domains of antibodies, nanobodies, and T-cell receptors. Its central claim is that explicitly labeling each training structure as apo (unbound) or holo (bound) lets the model learn two distinct, physically meaningful conformations from a single amino-acid sequence, so that at inference a user can ask for either state directly. The paper reports state-of-the-art or competitive accuracy across public benchmarks and on a private set of 286 novel antibodies with unseen CDR H3 loops, while running in under a second per variable domain, far faster than the diffusion-based general predictors it is compared against. If right, this gives therapeutic design a fast way to generate both the unbound and antigen-bound shape of an immune receptor, which matters for docking and for antibodies whose binding follows an induced-fit mechanism.","feed_headline":"Ibex predicts both bound and unbound antibody structures","feed_subtitle":"Ibex outputs two accurate structures from one sequence, under a second per domain.","key_machinery":"The load-bearing component is the conformation token: a one-hot input feature labeling each structure as apo or holo during training, which can be set at inference to request a bound or unbound prediction. Around it, Ibex uses 16 AlphaFold2-style structure module blocks with invariant point attention, a residual connection from the initial embedding to every structure module to preserve the token's influence, ESM-C 300M sequence embeddings, and a three-stage curriculum loss (FAPE, torsion, pLDDT, then structural violation losses) over labeled experimental and distilled data. The final model is an ensemble of eight models returning the prediction closest to the mean.","core_discovery":"The paper's central discovery is that a structure prediction model can be conditioned on binding state by means of a single conformation token. Trained on structures labeled apo or holo, Ibex produces two accurate, distinct structures from one sequence, recapitulating observed apo/holo conformational differences on 562 matched pairs and predicting hydrogen-bond networks characteristic of each state. The authors argue that previous models, trained on undifferentiated structural databases with multiple entries per sequence, risk predicting a non-physical average or collapsing to the most common state; the conformation token resolves this ambiguity. On the ImmuneBuilder test set Ibex reaches the lowest mean RMSD on TCR CDR beta3 and alpha3 loops and the second lowest on antibody and nanobody CDR H3 loops, and on the private out-of-distribution dataset of 286 novel antibodies it achieves the lowest mean CDR H3 RMSD (2.28 Å) among all compared methods. The authors attribute the out-of-distribution robustness to a three-stage curriculum over experimental immune structures, immunoglobulin-like domains, and a 60k-structure distillation set from predicted models.","pith_inferences":["The conformation token is a general architectural trick: since the network already learns to separate conformational states in a single latent space, the same token could be repurposed to condition on other discrete biological states (pH, allosteric ligand binding, oxidation state) without changing the model core, so long as labeled structures exist for those states; the paper suggests the idea bu","The apo/holo labeling rule (metadata-only) likely injects label noise: a crystal structure solved without its antigen in the asymmetric unit, or a non-physiological crystal contact, will be mislabeled. The paper's own ablation does not include a label-noise robustness test, so a direct testable extension is to train a version with a fraction of shuffled apo/holo labels and measure how much of the ","Because the private benchmark set comes from the same industrial pipeline that provided the high-resolution structures, the out-of-distribution numbers are not independently reproducible by outside groups; an independent check on publicly deposited, recently released apo/holo pairs would establish whether the generalization claim holds beyond the training distribution.","If Ibex's two predicted states are faithful, the difference vector between apo and holo predictions could serve as a cheap prior for flexible-docking or induced-fit studies, providing starting conformations for refinement without running MD; the paper does not report any docking experiments."],"forward_implications":["Two structures from one sequence: for any immune receptor sequence, Ibex can output both the apo and holo conformation, so designers can pick the relevant starting state for docking rather than using a single averaged model.","State-of-the-art out-of-distribution accuracy on novel CDR H3 loops: on the private dataset of 286 antibodies whose H3 loops differ from all public structures, Ibex's mean CDR H3 RMSD is 2.28 Å, beating both specialized tools like ABodyBuilder3 and general models like Boltz-1 and Boltz-2, suggesting it generalizes to the novel sequences that arise in real therapeutic programs.","The model's hydrogen-bond network predictions in the CDRs match the apo/holo ground-truth states in magnitude and connectivity, suggesting the two predicted conformations carry biophysical meaning beyond backbone geometry, with a slight bias toward overproducing holo-state hydrogen bonds.","Diffusion-based general predictors (Boltz-1, Chai-1) show almost no improvement in CDR H3 loop accuracy when sampled up to 1000 seeds, whereas Ibex directly conditions on the target state; this argues that for this loop, stochastic sampling does not substitute for explicit binding-state conditioning.","Ibex runs in 0.7 s on a single A10G GPU for a variable domain, roughly 10x faster than ESMFold and 90x faster than Boltz-2 including MSA, making high-throughput modeling of immune repertoires practical."],"supporting_citations":[{"why":"Supplies the AlphaFold2 structure module blocks and FAPE loss that Ibex adapts.","marker":"Jumper et al., 2021"},{"why":"Supplies SAbDab, the antibody structure database from which apo/holo labeled experimental structures are curated.","marker":"Dunbar et al., 2014"},{"why":"Supplies STCRDab, the source of labeled TCR structures.","marker":"Leem et al., 2018"},{"why":"Provides the ImmuneBuilder test set and the ABodyBuilder2 architecture that Ibex extends.","marker":"Abanades et al., 2023"},{"why":"Provides ABodyBuilder3, both a baseline for comparison and the codebase Ibex builds upon.","marker":"Kenlay et al., 2024"},{"why":"Provides ESMFold, used to generate the distillation set and as a baseline.","marker":"Lin et al., 2023"},{"why":"Provides Boltz-1, used both for distillation and as a baseline.","marker":"Wohlwend et al., 2024"},{"why":"Provides the Rosetta all-atom energy function used to validate hydrogen-bond networks in apo/holo predictions.","marker":"Alford et al., 2017"},{"why":"Provides Chai-1, a baseline for benchmarks and the seed-sampling experiment.","marker":"team et al., 2024"}],"fun_headline_variants":["Ibex predicts bound and unbound shapes from one sequence","Binding-aware AI models antibody, nanobody, TCR structures","One sequence, two accurate immune structures from Ibex","Ibex: fast, accurate structure prediction for immune proteins","Ibex: two immune structures per sequence in under a second"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire apo/holo distinction hangs on the metadata rule that a structure is 'apo' when no antigen chain appears in SAbDab or STCRDab, and 'holo' otherwise; if that binary labeling is noisy, the conformation token learns to separate artifacts rather than true bound and unbound states.","fun_headline_variants_meta":{"raw":{"variants":["Ibex predicts bound and unbound shapes from one sequence","Binding-aware AI models antibody, nanobody, TCR structures","One sequence, two accurate immune structures from Ibex","Ibex: fast, accurate structure prediction for immune proteins","Ibex: two immune structures per sequence in under a second"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000824,"raw_usage":{"total_tokens":3561,"prompt_tokens":860,"completion_tokens":2701,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":2630}},"tokens_in":476,"tokens_out":2701,"duration_ms":23785,"temperature":1.0,"reasoning_tokens":2630,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:05:44.375740+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of antibody–antigen complexes crystallized both with and without the antigen under identical conditions, ideally newly deposited ones absent from training; if Ibex's apo and holo predictions do not separate into the experimentally observed states, or if shuffling the apo/holo labels during training leaves the two-state prediction gap unchanged, the binding-state conditioning is not capturing real conformational differences.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies SAbDab, the antibody structure database from which apo/holo labeled experimental structures are curated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies STCRDab, the source of labeled TCR structures."},{"cited_title":"K., Boyles, F., Georges, G., Bujotzek, A., and Deane, C","cited_arxiv_id":null,"evidence_quote":"Provides the ImmuneBuilder test set and the ABodyBuilder2 architecture that Ibex extends."},{"cited_title":"A., Cutting, D., Nissley, D., and Deane, C","cited_arxiv_id":null,"evidence_quote":"Provides ABodyBuilder3, both a baseline for comparison and the codebase Ibex builds upon."},{"cited_title":"F., Leaver-Fay, A., Jeliazkov, J","cited_arxiv_id":null,"evidence_quote":"Provides the Rosetta all-atom energy function used to validate hydrogen-bond networks in apo/holo predictions."}],"review_version":1}