{"id":"ee1a9cf2-7ad1-433c-a7fc-61d7d1136d5f","arxiv_id":"2411.13390","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Energy-based sampling with a human-antibody prior generates diverse heavy-chain mutants near the wild type that lie on predicted affinity-solubility Pareto fronts and outperform constrained local search in synthetic benchmarks.","lead":"This paper trains energy-based generative models to propose antibody mutations that trade off predicted solubility, affinity, and human-likeness. It shows the method can trace a predicted Pareto front and beats a local-search baseline on synthetic antibody fitness landscapes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Solubility surrogate is the unvalidated link: generated Pareto candidates reach fsol scores 2-6, far beyond the 83 validated antibodies (mean ~0.7, Spearman 0.40), and the synthetic benchmark reuses the same surrogate, so real solubility gains are unverified.","rationale":"The paper's methodological contribution, energy-based sampling from a human-sequence prior to optimize a multi-objective surrogate, is credible and is supported on synthetic affinity ground-truth tasks in Section II.E, where MCMC and GFlowNet often outperform the constrained local-search baseline. The reader's CONDITIONAL verdict is therefore appropriate. My stress-test refines the reader's weakest assumption: the two surrogate models are not equally unvalidated. The affinity surrogate has a synthetic ground-truth check and a direct validation (Pearson r=0.58 on held-out data); the solubility surrogate does not. The real-task Pareto front in Fig. 3 is plotted in the space of predicted fsol, and the generated values extend to 2-6, far beyond the 83 validated antibodies whose average is approximately 0.7. The synthetic task reuses the same predicted fsol as the objective, so it demonstrates only surrogate optimization for solubility, not correspondence to true solubility. This makes the solubility leg of the Pareto-optimal antibodies claim the load-bearing unverified step. A focused test, independent solubility scoring or a small HIC experiment on 4-6 mutation candidates, would settle it. If that test fails, the appropriate correction is to reframe the real-task claim as a predicted Pareto front or to add experimental validation, which is exactly the condition the reader attached. I therefore leave the verdict unchanged and mark partial agreement with the reader's weakest-assumption analysis.","tokens_in":19417,"tokens_out":6977,"duration_ms":82793,"concrete_test":"Take the top B=500 generated CB-119 candidates from Section II.D, score them with an independent solubility predictor not used in training, such as CamSol or a separately trained SASA/HIC model, and compute the Spearman correlation with the paper's fsol, stratified by Hamming distance (1-3 vs 4-6). If the correlation for 4-6 mutation sequences is not significantly positive, or if the generated candidates are not enriched over random 4-6 mutation mutants in the independent solubility score, the solubility leg of the Pareto claim fails. A decisive wet-lab check would be HIC retention time measurements on roughly 50 candidates spanning the predicted front; the claim survives only if measured HIC retention time correlates with predicted fsol on those candidates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The real-antibody Pareto claim in Section II.D depends on the solubility surrogate, and this link is unvalidated exactly where it matters. The solubility model (Eq. 4) achieves only Spearman r=0.40 against HIC retention time on 83 clinical antibodies, with those antibodies having an average predicted score of roughly 0.7. The generated candidates in Fig. 3A/B have predicted fsol values of roughly 2 to 6, a regime with zero validation; the paper itself notes this exceeds any therapeutic antibody used to validate the model. The synthetic benchmark in Section II.E cannot rescue this: its ground-truth affinity function is synthetic, but the 'solubility' objective in those experiments is still the same predicted fsol. Thus the synthetic results validate only that the sampler optimizes the surrogate, not that the surrogate measures real solubility. If fsol is biased or saturating for sequences with 4-6 mutations, the claimed Pareto front between solubility and affinity is an artifact of the proxy, and the high-solubility candidates need not be more soluble than the wild type. The affinity half of the claim is better supported because the synthetic task provides a true affinity ground truth; the solubility half has no analogous check.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an energy-based generative approach for single-round optimization of monoclonal antibody heavy-chain sequences. The generative distribution is p(x) ∝ pHUM(x) exp(-E(x)/T), with pHUM a human-antibody language model prior and E a linear combination of predicted affinity and predicted solubility. The authors implement sampling with Metropolis-Hastings and GFlowNet, restrict mutations to at most six positions from the wild-type CB-119 binder AB-14, and select candidates near an empirical Pareto front defined by the two predictions. The real-sequence part is complemented by a synthetic benchmark with an epistatic ground-truth affinity function, where the method is compared with antBO local search, a random baseline, and the training set. The central methodological claims are that the samplers produce diverse, novel, and Pareto-approximate sequences, and that the energy-based approach outperforms constrained local search on the synthetic tasks.","tokens_in":19694,"tokens_out":4634,"duration_ms":51339,"significance":"If the method's outputs are interpreted as candidates for wet-lab screening within a larger pipeline, the paper makes a useful contribution: it clearly demonstrates that energy-based sampling with a human-sequence prior can navigate a restricted mutation space and find diverse sequences that optimize in silico objectives, and the synthetic benchmark with known affinity ground truth provides a meaningful internal check of the optimization procedure. The manuscript is also strong in transparency: the Boltzmann derivation, the MCMC and GFlowNet implementations, and the comparison protocols are described in enough detail to be reproduced, and code and data are promised at a public repository. The main limitation is external validity: the real CB-119 Pareto front is built entirely from proxy predictions, and the solubility surrogate is validated only in a low-score regime, so the paper's more ambitious framing of generating antibodies that are actually Pareto-optimal in affinity and solubility is not supported by the evidence presented.","major_comments":[{"comment":"The solubility surrogate is the main unvalidated link in the real-sequence Pareto claim. The model achieves Spearman r = 0.40 on 83 clinical-stage antibodies whose mean predicted solubility score is approximately 0.7 (Section II.C), yet the generated candidates in Fig. 3A/B have predicted fsol values of roughly 2 to 6, a regime with no validation points. Because the empirical Pareto front in Section II.D is defined using this same fsol, the claim that the generated sequences lie on a Pareto front between affinity and solubility is only a statement about the surrogate, not about measured solubility. The synthetic benchmark in Section II.E cannot resolve this, since the 'solubility' objective there is the same predicted fsol; it validates that the samplers optimize the surrogate but not that the surrogate measures real solubility. I request that the authors either explicitly reframe all real-sequence Pareto statements as being with respect to predicted properties, or provide additional validation (e.g., experimental HIC measurements, or at least a calibration analysis showing that fsol extrapolates beyond the validated range).","section":"II.C, II.D, Eq. (4), Fig. 3"},{"comment":"The affinity model is validated only on the training distribution of single, double, and triple mutants, with Pearson r = 0.58 on held-out data, but the generative process explores sequences with up to six mutations from the wild type. The GP's uncertainty estimate σ(x) is used in the acquisition function (β values) but does not by itself correct for systematic extrapolation error at mutational distances 4–6. The paper should quantify how many generated sequences fall at Hamming distance 4, 5, and 6 from the training set, report the associated GP predictive variances, and ideally use the synthetic benchmark to show that the method's advantage over antBO persists when the GP is trained only on up-to-triple mutants but evaluated on 4–6 mutation ground truth. Without this, the single-round optimization claim rests on an untested extrapolation.","section":"II.B and II.D"},{"comment":"The synthetic benchmark provides a genuine ground-truth for affinity, but the solubility objective in the synthetic tasks is the very same predicted fsol used in the real-task energy. Consequently, the comparison with antBO demonstrates superiority at optimizing a composite surrogate but does not provide evidence about real solubility. This should be stated explicitly in the text, and the conclusion 'our energy based sampling method performs better than constrained optimization' (end of Section II.E) should be restricted to the in silico objectives. A cleaner test would use a second, independent solubility proxy, or a synthetic ground-truth solubility model, in the benchmark.","section":"II.E, Figs. 5–6"},{"comment":"There is a sign inconsistency in the derivation of the Boltzmann distribution. Equation (17) states p = arg max_π (Σ_x π(x) E(x) + T DKL(π||pHUM)), but the text immediately says 'the solution of this minimization is given by the Boltzmann law, Eq. 1.' Maximizing Σ π E + T DKL would favor high-energy, high-entropy distributions, not the low-energy distribution exp(-E/T)/Z. The correct formulation is a minimization (or a maximization with -E and -T DKL). Since this equation is the theoretical foundation of the method, it should be corrected; implementation-wise the authors do use Eq. (1), so this appears fixable as a sign/optimization-direction error, but it must be addressed.","section":"IV.D, Eq. (17)"}],"minor_comments":[{"comment":"Typo: 'time time it takes' should read 'time it takes'.","section":"II.C"},{"comment":"The second term in the distance-to-Pareto-front expression is missing a superscript: it should read (fsol(x) − fsol(x'))² / σ²_sol. The same typo appears in the displayed equation.","section":"Eq. (5)"},{"comment":"The main text sets the solubility threshold to fmin_sol = 4.0, but the caption of Fig. 5B/D says 'predicted solubility score above 3.' Please align the threshold value between text and caption.","section":"II.E, Fig. 5 caption"},{"comment":"Panel B of Fig. S5 is described as 'on the simple synthetic task' but the figure shows the hard epistasis model; this is likely a copy-paste error and should be corrected.","section":"Fig. S5 caption"},{"comment":"Typo: 'distance to the first amino acidi' should be 'distance to the first amino acid i'.","section":"IV.A"},{"comment":"Reference [33] has a placeholder 'year?' and should be completed before publication.","section":"References"},{"comment":"The notation hi(i) should be hi(xi), since the single-mutation effect depends on the amino acid identity, not just the position.","section":"Eq. (8)"},{"comment":"In the sentence introducing the inverse temperature, 'cloneness' appears to be a typo for 'closeness'.","section":"II.D"}],"recommendation":"major_revision","confidential_remarks":"The paper's methodological core—energy-based sampling with a human-sequence prior, MCMC and GFlowNet samplers, and a synthetic benchmark with true affinity ground truth—is sound and potentially useful for computational antibody design. The main concern is that the real-sequence Pareto front is constructed entirely from proxy predictions, and the solubility surrogate in particular is unvalidated exactly in the regime where the generated candidates lie. This is a correctable framing issue rather than a fatal flaw, but it requires substantial revisions to the claims and, ideally, additional analyses of extrapolation behavior. The authors' affiliation with Sanofi and the funder's involvement are disclosed; the manuscript appears to be an honest description of a computational method, and I do not see grounds for questioning its integrity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid methods paper, and the main caveat is exactly where you'd expect: the solubility surrogate is the weakest link, and the CB-119 Pareto front should be read as 'predicted Pareto front,' not validated biology. But the core machinery is honestly tested, and the paper earns a serious referee.\n\nWhat's new: KL-regularized Boltzmann sampling with a human heavy-chain language model prior, combined with a GP affinity predictor and a CNN-SASA solubility predictor; a systematic UCB/LCB acquisition sweep; and a careful synthetic benchmark against antBO and random mutants. The MCMC implementation is transparent, the GFlowNet comparison is honest (MCMC often wins), and the code and data are on GitHub. That's reproducible, formal work, and it deserves credit.\n\nThe synthetic validation is the strongest part: with a true epistatic affinity function as ground truth, the method finds more high-affinity sequences than antBO in most settings, and the random baseline beats everyone on the easy task—which they honestly report. That tells you the method is a real optimizer, not a vacuous one.\n\nThe soft spots, in proportion:\n\n- The solubility surrogate is validated at Spearman 0.40 on 83 antibodies with scores around 0.7, and the generated candidates have predicted fsol of 2 to 6—entirely outside that range. The synthetic benchmark reuses the same surrogate for the 'solubility' objective, so it validates that the sampler optimizes the surrogate, not that the surrogate measures real solubility. This is the load-bearing caveat. The paper admits it in the Discussion, but the abstract and Section II.D state the Pareto front without that qualifier. Reframe as predicted Pareto front.\n\n- The GP affinity model has Pearson 0.58 on up to triple mutants, but generated sequences go to 6 mutations. The synthetic experiments on a true affinity function help, but they use the same type of GP; extrapolation to 4–6 mutations in the real task remains unverified.\n\n- Several hyperparameters (T, beta, fmin_sol, dlim) are hand-tuned. They do explore T and beta sensitivity, which is decent, but no principled selection.\n\nNone of this sinks the paper. The method is credible, the code is available, and the limitations are openly stated. For a reader in antibody design or ML-guided protein engineering, this is worth reading and citing. It deserves peer review; I'd recommend acceptance with revisions that explicitly scope the CB-119 claims as predicted, add uncertainty bars or a discussion of the extrapolation risk, and pin a commit hash to the repository.\n\nBring it to reading group—it's a good case study in how to do synthetic benchmarks honestly, and where proxy validation can still mislead.","headline":"A solid methods paper with honest synthetic benchmarks; the CB-119 Pareto front is a predicted front, and the solubility surrogate is the one unvalidated link.","tokens_in":20257,"tokens_out":3492,"would_cite":true,"duration_ms":36441,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that sampling from a human-antibody distribution biased by predicted affinity and solubility generates diverse, novel heavy-chain mutants along the Pareto front.","keywords":["energy-based generative models","monoclonal antibody optimization","Pareto front","affinity prediction","solubility prediction","GFlowNet","Metropolis-Hastings sampling","human antibody sequence prior"],"falsifier":"Measure true dissociation constants (e.g., by surface plasmon resonance or AlphaSeq) and true HIC retention times for a few hundred of the generated four-to-six-mutation heavy chains, and compare the resulting empirical Pareto front with the predicted one; if the correlation between predicted and measured affinity on these held-out mutants is close to zero, or if the measured front is no better than that of random six-mutation mutants, the central claim fails.","tokens_in":19179,"feed_emoji":"🧬","tokens_out":11554,"duration_ms":109375,"temperature":0.7,"pith_summary":"The paper seeks to establish that energy-based generative sampling can perform single-round multi-objective optimization of an antibody heavy chain. Starting from a wild-type binder, the model draws candidates from a Boltzmann distribution that stays close to the distribution of human antibody sequences while favoring high predicted affinity and high predicted solubility, with mutations capped at six positions. The authors show that the sampled heavy chains form a continuous empirical Pareto front in the two predicted properties, with diversity and novelty far from the training data. They then confirm on synthetic epistatic landscapes with known ground truth that the generated candidates beat a constrained local-search baseline, and that an optimistic acquisition policy works best unless the test budget is very small.","feed_headline":"Antibody design by energy-based sampling hits the Pareto front","feed_subtitle":"A generative sampler balances human likeness, affinity, and solubility, and beats constrained local search on synthetic benchmarks.","key_machinery":"The load-bearing object is the Boltzmann distribution over heavy-chain sequences, $p(x) = \\frac{1}{Z} p_{\\mathrm{HUM}}(x) \\exp(-E(x)/T)$, where $p_{\\mathrm{HUM}}$ is an autoregressive transformer trained on human heavy chains, $E(x) = -w \\hat{f}_{\\mathrm{aff}}(x) - (1-w)\\hat{f}_{\\mathrm{sol}}(x)$ combines a Gaussian-process affinity predictor and a SASA-plus-hydrophobicity solubility predictor, $T$ controls how far samples may wander from the human-antibody prior, and $w$ moves weight between affinity and solubility. The acquisition parameter $\\beta$ enters through $\\hat{f}_{\\mathrm{aff}}(x) = \\mu(x) + \\beta \\sigma(x)$, letting the sampler be pessimistic or optimistic about prediction uncertainty. Two samplers are used: Metropolis-Hastings over a single-mutation neighborhood with a six-mutation cap, and GFlowNet, an amortized generative model. The final selection step ranks samples by their distance to the empirical Pareto front.","core_discovery":"On the paper's own terms, the discovery is that the distribution $p(x) \\propto p_{\\mathrm{HUM}}(x)\\,e^{-E(x)/T}$, with $E(x) = -w \\hat{f}_{\\mathrm{aff}}(x) - (1-w)\\hat{f}_{\\mathrm{sol}}(x)$, is a practical generative model for lead optimization. Sampling this distribution by Metropolis-Hastings or GFlowNet produces heavy-chain mutants within six mutations of the wild type that lie on the empirical Pareto frontier between predicted affinity and predicted solubility; higher solubility can be obtained at the cost of roughly an order of magnitude in predicted affinity. Because the samples are weighted by $p_{\\mathrm{HUM}}$, they resemble natural human antibodies, and the temperature $T$ sets the diversity-versus-optimality trade-off. In synthetic tasks where the epistatic affinity function is known exactly, the same procedure generates more sequences above fixed affinity and solubility thresholds than the constrained local search baseline, and an optimistic choice of the acquisition parameter $\\beta$ (using the upper confidence bound) is generally best. The paper explicitly states that the method's purpose is to 'generate diverse and novel sequences along the Pareto front between affinity and solubility.'","pith_inferences":["Because the affinity model's validation is limited to single-to-triple mutants (Pearson $r=0.58$) and the solubility model reaches Spearman $r=0.40$, the practical value of the generated Pareto front depends on whether these correlations persist at four to six mutations; a direct wet-lab check on generated sequences would settle this.","The same energy-based scheme should extend to more than two objectives, as the authors note, but linear scalarization with fixed weights can only sweep convex regions of the front; the temperature term may be essential for exploring non-convex trade-offs.","A solubility predictor with per-residue uncertainty estimates would let the acquisition function penalize or reward uncertain residues, potentially improving small-budget performance.","The synthetic results suggest that ruggedness of the fitness landscape, not just noise, determines the best optimism level: the harder the epistatic task, the larger the $\\beta$ needed, until the budget becomes tiny."],"forward_implications":["The generated heavy chains populate the predicted Pareto front continuously, so a developer can choose the trade-off between affinity and solubility rather than commit to a hard threshold.","Increasing the predicted solubility score from roughly 2 to 6 costs about one order of magnitude in predicted affinity, quantifying the trade-off for the CB-119 binder.","Sampled sets have mean pairwise Hamming distance between 4.5 and 5 (out of a maximum 12 under the six-mutation cap) and novelty between 2.5 and 4, so the method produces diverse, previously unseen candidates.","On synthetic tasks, an optimistic acquisition parameter ($\\beta = 1$ or $2$, depending on difficulty) outperforms a pessimistic one, except for very small budgets where conservative $\\beta = 0$ is better.","Both samplers beat the constrained local-search baseline on most synthetic tasks; GFlowNet only clearly wins on the hardest task with a solubility threshold."],"supporting_citations":[{"why":"Supplies the autoregressive human heavy-chain model used as the humanness prior $p_{\\mathrm{HUM}}$.","marker":"[11]"},{"why":"Supplies the CB-119/SARS-CoV-2 antibody affinity dataset used to train the Gaussian process.","marker":"[15]"},{"why":"Supplies the SASA-plus-hydrophobicity solubility score and its learned weights, the basis of $\\hat{f}_{\\mathrm{sol}}$.","marker":"[20]"},{"why":"Defines the constrained local-search baseline that the synthetic comparisons must beat.","marker":"[26]"},{"why":"Introduces the GFlowNet sequence-design method and the diversity and novelty metrics used for evaluation.","marker":"[9]"},{"why":"Provides the Gaussian-process implementation used for affinity prediction.","marker":"[16]"},{"why":"Provides the convolutional architecture repurposed for per-residue solvent-accessible-surface-area prediction.","marker":"[22]"},{"why":"Supplies the clinical-stage antibody HIC retention times used to validate the solubility predictor.","marker":"[1]"}],"fun_headline_variants":["Energy-based sampling puts antibody design on the Pareto front","Generative model finds Pareto-optimal antibody mutants","Energy-based generative models beat local search for antibody leads","Sampling antibodies on the Pareto frontier with energy-based models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole approach assumes the proxy models for affinity and solubility remain accurate for sequences with four to six mutations, even though the affinity model was validated on single, double, and triple mutants and the solubility model's correlation with measured retention time is moderate.","fun_headline_variants_meta":{"raw":{"variants":["Energy-based sampling puts antibody design on the Pareto front","Generative model finds Pareto-optimal antibody mutants","Energy-based generative models beat local search for antibody leads","Sampling antibodies on the Pareto frontier with energy-based models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000898,"raw_usage":{"total_tokens":3883,"prompt_tokens":976,"completion_tokens":2907,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":2844}},"tokens_in":592,"tokens_out":2907,"duration_ms":23786,"temperature":1.0,"reasoning_tokens":2844,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:28:15.537069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure true dissociation constants (e.g., by surface plasmon resonance or AlphaSeq) and true HIC retention times for a few hundred of the generated four-to-six-mutation heavy chains, and compare the resulting empirical Pareto front with the predicted one; if the correlation between predicted and measured affinity on these held-out mutants is close to zero, or if the measured front is no better than that of random six-mutation mutants, the central claim fails.","supporting_citations":[{"cited_title":"BioRxiv pp 2021– 12","cited_arxiv_id":null,"evidence_quote":"Supplies the autoregressive human heavy-chain model used as the humanness prior $p_{\\mathrm{HUM}}$."},{"cited_title":"(2022) A dataset comprised of binding interactions for 104,972 antibodies against a sars-cov-2 peptide","cited_arxiv_id":null,"evidence_quote":"Supplies the CB-119/SARS-CoV-2 antibody affinity dataset used to train the Gaussian process."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SASA-plus-hydrophobicity solubility score and its learned weights, the basis of $\\hat{f}_{\\mathrm{sol}}$."},{"cited_title":"(2014) Sabdab: the structural antibody database","cited_arxiv_id":null,"evidence_quote":"Defines the constrained local-search baseline that the synthetic comparisons must beat."},{"cited_title":"(2022) Biological sequence design with gflownets (PMLR), pp 9786–9801","cited_arxiv_id":null,"evidence_quote":"Introduces the GFlowNet sequence-design method and the diversity and novelty metrics used for evaluation."},{"cited_title":"Advances in neu- ral information processing systems 31","cited_arxiv_id":null,"evidence_quote":"Provides the Gaussian-process implementation used for affinity prediction."},{"cited_title":"Frontiers in immunology 13:958584","cited_arxiv_id":null,"evidence_quote":"Provides the convolutional architecture repurposed for per-residue solvent-accessible-surface-area prediction."},{"cited_title":"(2017) Biophysical properties of the clinical- stage antibody landscape","cited_arxiv_id":null,"evidence_quote":"Supplies the clinical-stage antibody HIC retention times used to validate the solubility predictor."}],"review_version":1}