{"id":"b861010c-5f1a-418c-a3e4-b8d00253d5fa","arxiv_id":"2607.20057","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"AAMFM combines ESM3, an antigen-geometry adapter, and Cal-DPO preference optimization rewarded by AlphaFold3-style scores to design antibody CDRs and structures, reporting higher predicted binding scores than prior methods.","lead":"The paper introduces AAMFM, an antibody-design model built on a large protein model (ESM3), adding antigen structure and epitope information through a small adapter and fine-tuning it with preference optimization scored by an AlphaFold3-style predictor. It reports higher predicted binding scores and better sequence plausibility than existing antibody co-design baselines, but its headline gains are measured on the same in-silico oracle used to train it.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Functional design claim rests on an unvalidated oracle that the model is trained to maximize; no independent binding metric supports the central claim.","rationale":"The reader's weakest_assumption identifies the exact load-bearing concern: the validity of AF3/Protenix confidence metrics as proxies for antibody functionality. This is more fundamental than the circular evaluation because if the proxy is invalid, the entire 'functional' claim collapses regardless of comparison fairness. The paper provides no experimental validation and the cited references for AF3-confidence affinity correlation are not sufficient to establish the proxy for de novo designed antibodies. I also note the training/evaluation oracle overlap compounds the problem: the model is optimized to maximize the same metrics used to demonstrate SOTA, so the comparisons to baselines are biased. The paper's own limitation statement confirms no in vitro validation. Therefore the reader's REJECT verdict is appropriate; no adjustment is needed, though a reframed claim about improving in-silico oracle scores could be acceptable in a revised version.","tokens_in":17148,"tokens_out":5703,"duration_ms":45886,"concrete_test":"Benchmark the Protenix AF3-score and ipTM against experimentally measured binding affinities across a standard antibody-antigen dataset (e.g., AB-Bind or SAbDab with measured KD). Compute the Spearman correlation between these confidence scores and log(KD). If ρ < 0.3, the AF3-score is not a valid functional proxy, undermining the claim. Alternatively, express 10–20 AAMFM-CalDPO and baseline designs and measure binding by SPR/BLI; if AAMFM designs do not show higher affinity or binding frequency, the functional claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract; §4.1) that AAMFM achieves state-of-the-art functional antibody design depends on the premise that Protenix/AlphaFold3 confidence metrics (AF3-score, ipTM) reliably indicate antibody-antigen binding functionality. This premise is not established. In §3.4, AF3-score is declared the 'primary oracle for antibody functionality' and used, with AntiBERTy PLL, to construct Cal-DPO preference labels. The same metrics are then the headline evaluation in §4.1 (Table 1) and are used to claim functional superiority. Because AAMFM-CalDPO is explicitly optimized to increase these scores, the evaluation is biased in its favor relative to baselines that were not optimized against this oracle. The paper admits (Conclusion) that no in vitro validation was performed. Prior work suggests AF3 confidence does not reliably predict binding affinity changes, so the high scores may reflect reward hacking rather than true function. Without an independent binding measurement, the 'functional antibody design' claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AAMFM, a multimodal antibody language model built on ESM3. An antigen geometric/epitope-aware adapter injects antigen structure and epitope information, followed by two fine-tuning stages: SFT on antibody-antigen complexes (SAbDab) and Cal-DPO using a preference dataset whose labels are derived from Protenix (an AlphaFold3 reproduction) AF3 scores and AntiBERTy pseudo-log-likelihood. Experiments on full-antibody co-design and CDR-H3 co-design compare against several baselines, reporting PLL, AF3-score, pTM, ipTM, AAR, RMSD and PHR. The main claim is state-of-the-art functional antibody design.","tokens_in":17285,"tokens_out":6592,"duration_ms":57265,"significance":"If the claims were well supported, the paper would be a meaningful contribution to computational antibody design: the multi-stage adaptation from a general protein foundation model, the lightweight antigen adapter, and the construction of a ~30k preference-pair dataset are all useful elements, and the code is promised to be public. However, the headline claim rests on an evaluation that is substantially circular: the training objective in §3.4 is the same family of metrics (Protenix AF3 score, AntiBERTy PLL) used as the primary evaluation axes in §4.1. Baselines were not trained against this oracle. Furthermore, no independent binding evidence is provided. Thus the current support for 'functional antibody design' is weak, although the methodological pieces could be salvaged with a major revision that adds independent validation or substantially reframes the claims.","major_comments":[{"comment":"The preference labels in §3.4 are defined by Protenix AF3 score and AntiBERTy PLL (dual thresholds 0.2 and 0.1). Table 1 then evaluates the model on exactly these metrics (PLL, AF3-score, pTM, ipTM). Since AAMFM-CalDPO is explicitly optimized to increase these scores, while the baselines are not, the reported gains largely measure reward optimization rather than functional superiority. This circularity is the central load-bearing issue. The authors should evaluate on an independent metric not used in training (e.g., experimentally measured binding affinities, or a held-out oracle) or clearly limit the claims to 'improving in-silico confidence scores'.","section":"§3.4, §4.1, Table 1"},{"comment":"The abstract and §4.1 claim state-of-the-art 'functional antibody design,' but the Conclusion admits that no in vitro validation was performed. AF3-score and ipTM are structure-prediction confidence metrics, not binding measurements. The paper itself notes the risk of reward hacking in §3.4, and the cited correlations (refs [35,21,62]) do not establish that maximizing these confidence scores yields functional antibodies. Without an independent binding assay, the term 'functional' is an overclaim. This is not merely a presentation issue; it undermines the central contribution.","section":"Abstract / Conclusion / §3.4"},{"comment":"The DPO preference dataset is generated from 'all PDB structures in the RAbD dataset' (Appendix A.3), while evaluation is reported on SAbDab (Tables 1–3). RAbD is derived from SAbDab, so it is unclear whether the complexes used for preference-data construction overlap with the evaluation set. If such overlap exists, the evaluation is further confounded by data leakage. The authors should specify the exact split and show that both SFT and DPO training data are disjoint from the test set.","section":"Appendix A.3, §4.1"},{"comment":"Two evaluation-fairness issues: (a) AbX is evaluated with provided checkpoints rather than being retrained on the same dataset as the other baselines (DiffAb, dyMEAN, GeoAB are retrained), which could bias the comparison. (b) No error bars or significance tests are reported. With only 10 samples per antigen and no variance information, small differences such as PLL -0.87 vs -0.91 in Table 1 may be within noise. Confidence intervals or paired tests should be provided to support the claimed superiority.","section":"Appendix A.1, Tables 1–3"}],"minor_comments":[{"comment":"Figure 1 is visually overloaded and the chain labels (e.g., 'H1FR1L2FR2L1FR1 L3FR3H3FR3H2FR2FR4FR4') are garbled and hard to read. A cleaner schematic would help the reader follow the pipeline.","section":"Figure 1"},{"comment":"Equation (11) writes the DPO loss as -E[log σ(β log πθ(yw)/πref(yw)) - log σ(β log πθ(yl)/πref(yl))], which is not the standard DPO loss. The main-text Eq. (3) is correct, so this is likely a typographical error; please fix it.","section":"Appendix D.2, Eq. (11)"},{"comment":"The dual thresholds (0.2 in AF3 score and 0.1 in PLL) are introduced without justification or sensitivity analysis. Please provide a rationale or a small study showing the results are robust to these choices.","section":"§3.4 'Preference Definition'"},{"comment":"The caption 'Ablation on CDRs design' is vague. Specify that this is the full six-CDR co-design setting consistent with §4.1, and clarify that the 'w/o DPO (SFT)' row is the same as AAMFM-SFT.","section":"Table 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The reader's rejection is defensible given the circular evaluation, but I see a path to revision: the authors could add an independent validation benchmark (e.g., known binding-affinity data) or substantially reframe the contribution as an in-silico confidence optimization pipeline. If no independent evidence is added and the 'functional' claim is retained, the paper would not be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on protein design or preference optimization. The paper does real engineering: ESM3 fine-tuned on OAS and SAbDab, a lightweight GearNet adapter injecting antigen geometry and epitope positions, and Cal-DPO with preferences from Protenix (an AF3 reproduction) plus AntiBERTy PLL. That specific combination is new, and the ablations are useful. The 30k-pair preference dataset is a concrete artifact that could be reused. But the central claim is overreaching. They call the AF3-score the 'primary oracle for antibody functionality,' use it to label preference pairs, then evaluate on the same family of metrics—AF3-score, ipTM, pTM, PLL. Baselines were not trained against this oracle, so the comparison is tilted. The paper openly admits no in vitro validation. AF3 confidence is not a validated proxy for binding affinity; prior work shows it correlates poorly with mutational effects on binding. So 'state-of-the-art functional antibody design' is not supported by the presented evidence. What is supported is a narrower claim: preference optimization improves predicted complex confidence and sequence plausibility, and the sequence/structure recovery metrics (AAR, RMSD) are competitive. On those independent metrics they are not uniformly SOTA, so the abstract oversells even that. Minor issues: no error bars or significance tests; AbX is evaluated with its own checkpoints while other baselines are retrained; and the model is given the epitope positions at inference, which presumes knowledge that the user may not have. These are fixable. I would take the paper seriously as a contribution to oracle-based preference optimization for antibody design, but the functional claim has to be reformulated or backed with real binding data. A referee should ask for either experimental validation (even a handful of SPR/ELISA hits) or evaluation on an independent affinity dataset. Without that, the right headline is 'predicted structural quality,' not 'functional.' If the authors reframe, the core pipeline is publishable; as it stands, the central claim does not hold.","headline":"A well-engineered antibody design pipeline whose 'functional' claim is unsupported by its own evaluation loop, but the methodological core is worth a serious referee.","tokens_in":666,"tokens_out":778,"would_cite":false,"duration_ms":82458,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multimodal antibody foundation model generates complete antigen-specific antibody designs conditioned on antigen geometry and epitope positions, and preference optimization against structural-confidence scores yields state-of-the-art in s","keywords":["antibody design","antigen-specific","multimodal protein language model","complementarity-determining regions","preference optimization","epitope conditioning","structure prediction confidence","CDR co-design"],"falsifier":"Collect a panel of 50–100 AAMFM-CalDPO designs spanning a range of predicted AF3 scores for one or two antigens, synthesize and express them, and measure binding by surface plasmon resonance or yeast display; if the rank order of measured affinities does not track the predicted AF3 score and ipTM, the functional-design claim is refuted even if all in silico metrics replicate.","tokens_in":16913,"feed_emoji":"🧬","tokens_out":6270,"duration_ms":56583,"temperature":0.7,"pith_summary":"The paper tries to show that a single multimodal antibody model can take an antigen structure, a specified epitope, and an antibody framework sequence, and jointly output the sequences of all six CDR loops plus the complete antibody structure. To do this it adapts a large general protein language model to antibody sequence–structure pairs, then injects antigen geometry and epitope information through a small cross-modal adapter. A final preference-optimization stage steers the model toward designs that score highly on structure-prediction confidence and on natural-antibody sequence plausibility. The paper reports that this pipeline, called AAMFM with Cal-DPO, outperforms previous antibody co-design methods on all four headline metrics (PLL, AF3 score, pTM, and ipTM) and on most CDR-level metrics. A sympathetic reader would care because it is a concrete route from an antigen's structure to plausible, antigen-specific antibody candidates without experimental screening.","feed_headline":"Antibody model writes all six CDRs from antigen geometry","feed_subtitle":"It turns antigen shape and epitope positions into complete antibody sequences and structures in one pass.","key_machinery":"The load-bearing mechanism is the combination of the Antigen Geometric-Epitope-aware Adapter and the Cal-DPO preference-alignment stage. The adapter is a lightweight residual module that conditions the base model on antigen geometry (via a graph neural network) and epitope positions (via an interface encoder), letting the model attend to the binding site during generation. Cal-DPO is a preference-optimization loss that, unlike plain DPO, also pulls the absolute log-likelihood ratios of preferred and dispreferred CDR sequences toward a fixed margin, so the model's confidence tracks the oracle signal instead of just ranking pairs. The oracle signal itself is the AF3 ranking/confidence score fr","core_discovery":"AAMFM's central claim is that unified representation learning over antibody sequence, antibody structure, and antigen context—rather than sequence-only language modeling or graph-only co-design—is what makes antigen-specific functional design work. The model is built on a multimodal protein language model, post-trained on roughly 1.4 million paired antibody sequences and structures, then fine-tuned on experimental antibody–antigen complexes. The Antigen Geometric-Epitope-aware Adapter, only about 5.1 million parameters, fuses graph-extracted antigen geometry and epitope-interface embeddings into the model's latent space via residual adjustment. The Cal-DPO stage then optimizes against a pref","pith_inferences":["If AF3 confidence scores and ipTM correlate strongly with real binding—which the paper assumes but does not test—then AAMFM's preference pipeline is a high-throughput in silico prefilter that could substantially shrink the sequence space needing wet-lab screening.","The adapter-plus-preference recipe is not antibody-specific in principle; the same conditioning on partner geometry and interface labels could be applied to nanobodies, peptide binders, or other protein–protein interaction design tasks.","Because the preference labels currently come entirely from structural-prediction confidence, a natural and testable extension is to rebuild the preference dataset with experimentally measured affinities for a small set of antigens and check whether Cal-DPO then improves real binding rates rather than only predicted ones.","The dual-threshold labeling rule (simultaneous margins on AF3 and PLL) could be tuned or replaced with other quality signals—solubility, developability, immunogenicity—to steer generation toward a different definition of 'functional'."],"forward_implications":["Given an antigen structure and epitope mask, AAMFM-CalDPO can produce full CDR sequences and antibody structure in one generation pass, removing the need for separate sequence and structure models.","The Cal-DPO stage intentionally trades some native-sequence recovery (AAR) for higher predicted binding confidence and plausibility, so users optimizing for binding can use the preference-aligned variant while users optimizing for human-likeness can use the SFT variant.","Ablations show that removing the antigen adapter causes the largest drop in functional metrics, and disabling epitope input degrades foldability and binding confidence—so antigen geometry and epitope conditioning, not just the language-model prior, carry the antigen-specific signal.","The released preference dataset of roughly 30,000 pairs with AF3 and PLL labels is a reusable resource for training or aligning other antibody design models.","For CDR-H3-only redesign, the same model beats dedicated CDR-H3 co-design baselines on predicted-binding metrics, suggesting the framework applies at both full-antibody and single-loop granularity."],"fun_headline_variants":["Antibody AI designs all CDRs from antigen geometry in one pass","AAMFM: antigen-aware antibody design with sequence and structure","Antigen context in, functional antibody out: AAMFM","One model, both antibody and antigen, designs functional CDRs","AAMFM generates antibody sequences from antigen geometry and epitopes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim collapses if Protenix/AF3 confidence scores and ipTM are only weakly correlated with true antibody–antigen binding: the whole preference signal and the headline evaluation metrics are built on those structural-confidence scores, and the paper explicitly notes that no in vitro validation of the designed antibodies was performed.","fun_headline_variants_meta":{"raw":{"variants":["Antibody AI designs all CDRs from antigen geometry in one pass","AAMFM: antigen-aware antibody design with sequence and structure","Antigen context in, functional antibody out: AAMFM","One model, both antibody and antigen, designs functional CDRs","AAMFM generates antibody sequences from antigen geometry and epitopes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001227,"raw_usage":{"total_tokens":4865,"prompt_tokens":713,"completion_tokens":4152,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":4064}},"tokens_in":457,"tokens_out":4152,"duration_ms":25178,"temperature":1.0,"reasoning_tokens":4064,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:53:50.682573+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a panel of 50–100 AAMFM-CalDPO designs spanning a range of predicted AF3 scores for one or two antigens, synthesize and express them, and measure binding by surface plasmon resonance or yeast display; if the rank order of measured affinities does not track the predicted AF3 score and ipTM, the functional-design claim is refuted even if all in silico metrics replicate.","supporting_citations":[],"review_version":1}