Pith. sign in

REVIEW 4 major objections 3 minor 49 references

GROOT: Effective Design of Biological Sequences with Limited Experimental Data

T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A latent graph smoother turns scarce labels into reliable protein optimizers.

desk verdict A sensible, incremental LSO method whose empirical gains on scarce-label protein tasks are real, but whose central pseudo-label-fidelity assumption is never directly tested. read the letter →

arxiv 2411.11265 v1 pith:SOUWMRWM submitted 2024-11-18 cs.LG q-bio.QM

classification cs.LGq-bio.QM
keywords latentspaceoptimizationproteindesignlabelpropagationlimitedlabeleddatasurrogatemodelfitnesslandscapesmoothingAAVcapsidgreenfluorescent
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that scarce labeled data does not have to cripple latent space optimization for biological sequences. The proposed method, GROOT, samples synthetic latent points by interpolating training embeddings with Gaussian noise, then smooths their fitness values with label propagation over a k-nearest-neighbor graph before training the surrogate model. The paper argues this densifies the fitness landscape so that gradient-based optimizers can find high-fitness regions instead of getting stuck near false negatives. Empirically, GROOT reports state-of-the-art fitness on AAV and GFP across the harder1–harder3 scarce-label benchmarks, including a 6-fold fitness improvement over the training set in GFP and 1.3x in AAV, while staying competitive on exact-oracle Design-Bench tasks. The reason to care is that wet-lab evaluation is costly, and the method needs neither an oracle during optimization nor vast labeled datasets.

What carries the argument

The central object is the latent kNN graph with interpolated nodes and label propagation. GROOT repeatedly samples a training embedding $x$ uniformly, draws noise $\epsilon \sim \mathcal{N}(0, I_d)$, and creates a synthetic node $z = \beta x + (1-\beta)\epsilon$ until the graph has $N$ nodes; edges are $k$-nearest neighbors under Euclidean distance, and fitness labels are updated over $m$ layers by $Y' = \alpha D^{-1/2}AD^{-1/2}Y + (1-\alpha)Y$. The $z$-formula is the exploration mechanism, label propagation is the pseudo-labeling mechanism, and the bound $2(1-\beta)\sqrt{d}$ from Proposition 2 is the reliability guarantee that keeps synthetic nodes inside a 'reliable zone' near the training hull.

What would settle it

Measure the correlation between Euclidean distance in the latent space and absolute fitness difference on held-out pairs from a protein family; if the correlation is near zero, the kNN label-propagation premise fails. A direct test: train GROOT on a family with a known one-mutation fitness cliff, and check whether the optimizer proposes the low-fitness mutant because a synthetic node near it received a high smoothed label.

Watch

Extended reading notes

Core claim

GROOT claims that a surrogate trained on pseudo-labeled synthetic latent points, generated by $z = \beta x + (1-\beta)\epsilon$ with $x$ a training embedding and $\epsilon \sim \mathcal{N}(0, I_d)$, can extrapolate beyond the training set while remaining reliable. The paper proves that, under its two assumptions, the synthetic points fall outside the training convex hull with probability tending to 1 as $d$ grows, and that their expected distance to the hull is bounded by $2(1-\beta)\sqrt{d}$. It then shows empirically that the smoothed surrogate, fitted to these labels with a simple MLP, enables gradient ascent and L-BFGS to beat previous baselines in the extreme-label regime of AAV and GFP, and that the same smoothing also improves a ReLSO surrogate.

Load-bearing premise

The method assumes that Euclidean distance in the ESM-2/VAE latent space tracks fitness similarity, so kNN label propagation over synthetic noisy nodes yields meaningful pseudo-labels.

Editorial extensions

If this is right

  • With GROOT's smoothing, a two-layer MLP surrogate plus gradient ascent or L-BFGS outperforms AdaLead, CbAS, Bayesian optimization, GFN-AL, PEX, GGS, and ReLSO on the AAV harder1–harder3 and GFP harder1–harder3 benchmarks.
  • Smoothing cuts surrogate error on train and holdout sets alike; on AAV harder1 the train MAE drops from 4.94 to 1.02 and the holdout MAE from 8.93 to 5.78.
  • When the labeled subset is randomly sampled, 20% of the harder3 data, under 100 sequences, already matches the best point of the full harder3 set; in the lowest-fitness subsampling, roughly 50%, about 200 sequences, is enough.
  • Adding the same smoothing to ReLSO raises its fitness by 60%, 64.7%, and 22.7% on AAV harder1, harder2, and harder3, indicating the mechanism transfers across surrogate architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If latent Euclidean distance is a good fitness-similarity proxy, the smoothing step could be attached to any pretrained protein language model encoder, not just the VAE used here; the paper does not test that substitution.
  • The theory makes $\beta$ the dial between exploration and reliability, so tuning $\beta$ per protein family could further improve extrapolation; the paper leaves $\beta$ fixed and mentions mutation-effect analysis as future work.
  • Because the encoder and surrogate are treated as interchangeable, the same graph-smoothing recipe should apply to RNA and small-molecule design tasks beyond the three Design-Bench domains tested.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper addresses latent space optimization (LSO) for biological sequence design under extreme label scarcity. GROOT builds a kNN graph over ESM-2/VAE latent embeddings of the labeled training sequences, augments the graph with synthetic latent nodes formed by interpolating training latents with Gaussian noise, assigns these synthetic nodes pseudo-labels via one round of label propagation, trains a shallow MLP surrogate on the enlarged node set, and then runs gradient-based optimization (gradient ascent or L-BFGS) in the latent space. The authors report results on AAV and GFP fitness optimization at several difficulty levels and on three Design-Bench tasks, and they provide two propositions bounding the probability that synthetic nodes fall outside the training convex hull and the expected distance of such nodes to that hull. The paper claims state-of-the-art results across all difficulties on the two protein benchmarks and a 6x/1.3x fitness improvement over the training set in extreme low-data settings.

Significance. If the pseudo-labels assigned to the synthetic nodes are faithful, GROOT is a simple and computationally attractive data-augmentation recipe for the limited-label regime, and the paper's ablations (Tables 4 and 5) show that the smoothing step consistently reduces surrogate MAE and improves optimized fitness. The method is domain-agnostic in principle, and the authors evaluate on multiple benchmarks and release code, which are concrete strengths. However, the central mechanism -- label propagation over synthetic latent nodes producing informative fitness labels -- is not directly validated, and there is a mismatch between the implemented interpolation formula and the theoretically analyzed one. The SOTA claim is also stated more strongly than Table 3 supports unless non-degenerate diversity is explicitly made part of the criterion. These issues are fixable but require additional evidence or a substantial rewriting of the claims.

major comments (4)
  1. [Algorithm 1, line 4; Propositions 1-2] Algorithm 1, line 4 defines z as −β*x + (1−β)*ε, while Propositions 1 and 2 and the surrounding text (Section 3.3, "interpolating the learned latent x with random noise") define z = β*x + (1−β)*ε. The sign is geometrically material: as β→1, the implemented formula sends z toward −x, which is not an interpolation near a training node and can lie far outside the training hull. Proposition 2's bound on D(z, Conv(X)) therefore does not apply to the nodes that the algorithm actually generates. Either the algorithm must be changed to match the theory or the theory must be re-derived for the implemented formula; as written, the "reliable zone" justification does not cover the method being evaluated.
  2. [Algorithm 2 / Eq. (4); Table 4] The load-bearing assumption that label propagation assigns meaningful fitness values to synthetic nodes is never tested. With N=20,000 graph nodes and |D|≤1,157, roughly 94-98% of the surrogate training set is synthetic; Algorithm 2 initializes those nodes to 0 and, with the Section 4.1 settings m=1 and α=0.2, updates them once via Eq. (4). Their pseudo-labels are small weighted averages heavily diluted by zero-valued synthetic neighbors. Table 4 shows that smoothing reduces the trained surrogate's MAE on training and holdout sets, but that measures smoothness of the regression function, not whether the pseudo-labels recover oracle fitness at synthetic nodes. The authors should decode a random sample of synthetic latent nodes and compare their pseudo-labels to oracle scores, or perform a held-out label-removal experiment, to establish the correlation on which the entire method depends.
  3. [Table 3; Section 1.1] Section 1.1 claims GROOT "achieves state-of-the-art results across all difficulties" on the AAV and GFP benchmarks, but Table 3 reports ReLSO fitness 0.94 on GFP harder1, harder2, and harder3, while GROOT obtains 0.88, 0.87, and 0.62, respectively. The footnote excludes ReLSO because its generated population has collapsed to a single sequence (zero diversity), which is a legitimate quality concern. However, the paper does not define the SOTA criterion that would exclude a high-fitness but degenerate population. If the claim is "best non-degenerate designs," that should be stated explicitly; on the raw fitness metric the unqualified SOTA assertion is not supported.
  4. [Section 3.4; Assumption 2] Propositions 1 and 2 rely on Assumption 2, that latent vectors are i.i.d. N(0,I_d). The VAE loss in Eq. (2) only encourages this via a KL term with weight η; it does not guarantee the assumption for finite data or for the ESM-2-initialized encoder used here. Since these propositions are presented as the theoretical justification for label propagation on synthetic nodes, the authors should provide at least a diagnostic of Gaussianity (or a sensitivity analysis varying η and the KL weight) for the actual latent embeddings used in the main experiments. The current statement that the assumption is "achieved" by the KL loss is too strong without such evidence.
minor comments (3)
  1. [Section 4.1 vs. Table 8] Section 4.1 states that the hyperparameters "are not finely tuned" and sets the number of propagation layers to N_layers = 1, while Table 8 reports per-task Optuna tuning with, for example, m = 4 for the GFP tasks. The text should clarify which hyperparameter settings produced the numbers in Table 3 and whether the main results use the Section 4.1 defaults or the tuned values from Table 8.
  2. [Figure 2] Figure 2 plots distance to the set of training nodes, whereas Proposition 2 concerns distance to the convex hull. Since the convex hull contains the training set, the plotted quantity is a conservative check, but the caption and text should state this relationship explicitly to avoid appearing to validate the proposition with a different quantity.
  3. [Algorithm 1, line 4] Apart from the sign issue, the notation "V← −V∪{z}" is unorthodox and should be written as "V ← V ∪ {z}" to avoid confusing a set update with a sign operation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: GROOT's pseudo-label smoothing is a data-augmentation step evaluated against external oracles, and its theoretical propositions are independent Gaussian-geometry statements.

full rationale

I walked the derivation chain: latent embeddings -> synthetic nodes z = beta*x + (1-beta)*eps -> kNN graph -> label propagation -> surrogate training -> gradient-based MBO. The pseudo-labels are generated from the training labels by label propagation, so they are augmented training targets, not first-principles predictions; the reported fitness gains are measured by external oracles (GFP/AAV oracles from Kirjner et al.; exact Design-Bench oracles), so the central empirical claim does not reduce to the method's own inputs. Propositions 1-2 are mathematical statements about N(0,I_d) vectors and convex hulls under stated Assumptions 1-2; they are not fitted to benchmark outcomes, and Section 4.3 independently checks the 100% outside-hull claim and the distance upper bound. The only self-citation ([42], Tran & Hy 2024) appears in a general list of directed-evolution methods and is not load-bearing. The main text's 'not finely tuned' statement conflicts with Appendix E's Optuna tuning per task, and the paper never directly validates that Euclidean latent distance tracks fitness; these are empirical overfitting/validation caveats, not demonstrated equivalences. The Limitations section likewise flags hyperparameter sensitivity, which I weigh as a practical limitation rather than circularity.

Assumptions & free parameters 8 free parameters · 4 assumptions · 0 invented entities

The central claim depends on a set of tuned hyperparameters (N, alpha, m, k), a fixed exploration coefficient beta, and two stated assumptions about the latent distribution and the small-sample regime. No new physical entities are introduced. The paper does not disclose the KL weight eta in the VAE loss, which is a minor omitted hyperparameter.

free parameters (8)
  • beta (interpolation coefficient) = not reported; 0.5 used in Proposition validation (Section 4.3)
    Controls exploration vs reliability; theoretical bounds scale with (1-beta); not included in Appendix E tuning ranges, so it is fixed by hand.
  • alpha (label propagation weight) = 0.2-0.6 per task (Table 8)
    Tuned via Optuna per task; affects how much smoothed labels are mixed with original labels.
  • N (number of graph nodes) = 4,000-20,000 per task (Table 8)
    Graph size for synthetic node generation; tuned via Optuna per task.
  • m (number of label propagation layers) = 1-6 per task (Table 8)
    Smoothing depth; tuned via Optuna per task.
  • k (number of nearest neighbors) = 4-8 per task (Table 8)
    kNN graph degree; tuned via Optuna per task.
  • gamma (weighted adjacency factor) = 1.0
    Appears in Equation (5); fixed by hand, not tuned.
  • d (latent dimension) = 320 for proteins, 128 for Design-Bench
    Architecture choice, not tuned; affects the theoretical bounds.
  • eta (VAE KL weight) = not reported
    Appears in Equation (2) and controls how close the latent is to N(0,I_d), which Assumption 2 depends on; its value is never specified.
assumptions (4)
  • domain assumption Assumption 1: N << exp(d/(2(C_beta^2+2)))
    States the labeled set is tiny compared to latent volume; used in Proposition 1.
  • domain assumption Assumption 2: latent vectors x are i.i.d. N(0,I_d)
    Used in Propositions 1 and 2; only approximately true because the VAE KL term penalizes deviation from the prior but does not guarantee exact Gaussianity of encoder outputs.
  • domain assumption Euclidean latent distance correlates with fitness
    Justifies kNN graph and label propagation; not directly measured in the paper.
  • ad hoc to paper Label propagation on the constructed kNN graph yields accurate pseudo-labels for synthetic points
    Core of the method; the paper provides only indirect evidence (lower MAE on holdout sets), not a direct validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GROOT: Effective Design of Biological Sequences with Limited Experimental Data." pith.science (2026). https://pith.science/paper/SOUWMRWM

@misc{pith2026241111265,
  author       = {Pith},
  title        = {Pith review of: GROOT: Effective Design of Biological Sequences with Limited Experimental Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SOUWMRWM}},
  note         = {Machine review of arXiv:2411.11265}
}
read the original abstract

Latent space optimization (LSO) is a powerful method for designing discrete, high-dimensional biological sequences that maximize expensive black-box functions, such as wet lab experiments. This is accomplished by learning a latent space from available data and using a surrogate model to guide optimization algorithms toward optimal outputs. However, existing methods struggle when labeled data is limited, as training the surrogate model with few labeled data points can lead to subpar outputs, offering no advantage over the training data itself. We address this challenge by introducing GROOT, a Graph-based Latent Smoothing for Biological Sequence Optimization. In particular, GROOT generates pseudo-labels for neighbors sampled around the training latent embeddings. These pseudo-labels are then refined and smoothed by Label Propagation. Additionally, we theoretically and empirically justify our approach, demonstrate GROOT's ability to extrapolate to regions beyond the training set while maintaining reliability within an upper bound of their expected distances from the training regions. We evaluate GROOT on various biological sequence design tasks, including protein optimization (GFP and AAV) and three tasks with exact oracles from Design-Bench. The results demonstrate that GROOT equalizes and surpasses existing methods without requiring access to black-box oracles or vast amounts of labeled data, highlighting its practicality and effectiveness. We release our code at https://anonymous.4open.science/r/GROOT-D554

Figures

Figures reproduced from arXiv: 2411.11265 by the authors.

Figure 1
Figure 1. Overall framework of GROOT. After encoding sequences into the latent space, we generate new samples by adding Gaussian noise to existing vectors. These synthetic data lie outside the training set’s convex hull but within a reliable zone, as their distances from the hull are below a certain upper bound. We construct a kNN graph and run label propagation to smooth and refine node labels. These nodes and their fitness … view at source ↗
Figure 2
Figure 2. Distance from generated nodes outside the convex [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. The performance of GROOT on AAV and GFP harder3 tasks when we vary the labeled data ratio 𝑟. The mean and standard deviation over 5 different runs are re￾ported. subsampled with the given ratio 𝑟, while in the lowest setting, we use the fraction 𝑟 of the dataset with the lowest fitness scores. Using the same hyperparameters, [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 34 canonical work pages

  1. [1]

    Michael Ahn, Henry Zhu, Kristian Hartikainen, Hugo Ponte, Abhishek Gupta, Sergey Levine, and Vikash Kumar. 2020. ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots. In Proceedings of the Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 100) , Leslie Pack Kael- bling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 1300...

  2. [2]

    Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19). As- sociation for Computing Machinery, New York, NY, USA, 2623–2631. https...

  3. [3]

    Frances H. Arnold. 1996. Directed evolution: Creating biocatalysts for the future. Chemical Engineering Science 51, 23 (1996), 5091–5102. https://doi.org/10.1016/ S0009-2509(96)00288-6

  4. [4]

    Frances H. Arnold. 2018. Directed Evolution: Bringing New Chemistry to Life. Angewandte Chemie International Edition 57, 16 (2018), 4143–4148. https://doi.org/10.1002/anie.201708408 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/anie.201708408

  5. [5]

    Randall Balestriero, Jerome Pesenti, and Yann LeCun. 2021. Learning in high dimension always amounts to extrapolation. arXiv preprint arXiv:2110.09485 (2021)

  6. [6]

    Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. OpenAI Gym. arXiv:1606.01540 [cs.LG] https://arxiv.org/abs/1606.01540

  7. [7]

    David Brookes, Hahnbeom Park, and Jennifer Listgarten. 2019. Conditioning by adaptive sampling for robust design. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 773–782. https://proceedings.mlr.press/v97/brookes19a.html

  8. [8]

    Brookes, Amirali Aghazadeh, and Jennifer Listgarten

    David H. Brookes, Amirali Aghazadeh, and Jennifer Listgarten. 2022. On the sparsity of fitness functions and implications for learning. Proceedings of the National Academy of Sciences 119, 1 (2022), e2109649118. https://doi.org/10.1073/ pnas.2109649118 arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.2109649118

Show all 49 references
  1. [9]

    David H Brookes, Amirali Aghazadeh, and Jennifer Listgarten. 2022. On the sparsity of fitness functions and implications for learning. Proceedings of the National Academy of Sciences 119, 1 (2022), e2109649118

  2. [10]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  3. [11]

    Bryant, Ali Bashir, Sam Sinai, Nina K

    Drew H. Bryant, Ali Bashir, Sam Sinai, Nina K. Jain, Pierce J. Ogden, Patrick F. Riley, George M. Church, Lucy J. Colwell, and Eric D. Kelsic. 2021. Deep diversi- fication of an AAV capsid protein by machine learning. Nature Biotechnology 39, 6 (01 Jun 2021), 691–696. https://...

  4. [12]

    Egbert Castro, Abhinav Godavarthi, Julian Rubinfien, Kevin Givechian, Dhanan- jay Bhaskar, and Smita Krishnaswamy. 2022. Transformer-based protein genera- tion with regularized latent space optimization. Nature Machine Intelligence 4, 10 (01 Oct 2022), 840–851. https://doi.org...

  5. [13]

    Venkat Chandrasekaran, Benjamin Recht, Pablo A Parrilo, and Alan S Willsky

  6. [14]

    Can Chen, Yingxueff Zhang, Jie Fu, Xue (Steve) Liu, and Mark Coates. 2022. Bidirectional Learning for Offline Infinite-width Model-based Optimization. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds....

  7. [15]

    Tianlai Chen, Pranay Vure, Rishab Pulugurta, and Pranam Chatterjee. 2023. AMP-Diffusion: Integrating Latent Diffusion with Protein Language Models for Antimicrobial Peptide Generation. In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop. https://openreview.net/forum?id...

  8. [16]

    Patrick Emami, Aidan Perreault, Jeffrey Law, David Biagioni, and Peter St. John

  9. [17]

    Fowler and Stanley Fields

    Douglas M. Fowler and Stanley Fields. 2014. Deep mutational scanning: a new style of protein science. Nature Methods 11, 8 (01 Aug 2014), 801–807. https: //doi.org/10.1038/nmeth.3027

  10. [18]

    Nathan C. Frey, Dan Berenberg, Karina Zadorozhny, Joseph Kleinhenz, Julien Lafrance-Vanasse, Isidro Hotzel, Yan Wu, Stephen Ra, Richard Bonneau, Kyunghyun Cho, Andreas Loukas, Vladimir Gligorijevic, and Saeed Saremi. 2024. Protein Discovery with Discrete Walk-Jump Sampling. In...

  11. [19]

    Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D

    Rafael Gómez-Bombarelli, Jennifer N. Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D. Hirzel, Ryan P. Adams, and Alán Aspuru- Guzik. 2018. Automatic Chemical Design Using a Data-Driven Conti...

  12. [20]

    Boyken, and David Baker

    Po-Ssu Huang, Scott E. Boyken, and David Baker. 2016. The coming of age of de novo protein design. Nature 537, 7620 (01 Sep 2016), 320–327. https: //doi.org/10.1038/nature19946

  13. [21]

    Qian Huang, Horace He, Abhay Singh, Ser-Nam Lim, and Austin Benson. 2021. Combining Label Propagation and Simple Models out-performs Graph Neural Networks. In International Conference on Learning Representations . https:// openreview.net/forum?id=8E1-f3VhX1o

  14. [22]

    Moksh Jain, Emmanuel Bengio, Alex Hernandez-Garcia, Jarrid Rector-Brooks, Bonaventure F. P. Dossou, Chanakya Ajit Ekbote, Jie Fu, Tianyu Zhang, Michael Kilgour, Dinghuai Zhang, Lena Simine, Payel Das, and Yoshua Bengio. 2022. Biological Sequence Design with GFlowNets. In Proce...

  15. [23]

    Kauffman and Edward D

    Stuart A. Kauffman and Edward D. Weinberger. 1989. The NK model of rugged fitness landscapes and its application to maturation of the immune response. Journal of Theoretical Biology 141, 2 (Nov. 1989), 211–245. https://doi.org/10. 1016/s0022-5193(89)80019-0

  16. [24]

    Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. 2023. Segment Anything. arXiv:2304.02643 (2023)

  17. [25]

    Jaakkola, Regina Barzilay, and Ila R Fiete

    Andrew Kirjner, Jason Yim, Raman Samusevich, Shahar Bracha, Tommi S. Jaakkola, Regina Barzilay, and Ila R Fiete. 2024. Improving protein optimization with smoothed fitness landscapes. In The Twelfth International Conference on Learning Representations. https://openreview.net/f...

  18. [26]

    Aviral Kumar and Sergey Levine. 2020. Model Inversion Networks for Model- Based Optimization. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 5126–5137. https://proc...

  19. [27]

    Minji Lee, Luiz Felipe Vecchietti, Hyunkyu Jung, Hyunjoo Ro, Meeyoung Cha, and Ho Min Kim. 2023. Protein Sequence Design in a Latent Space via Model-based Reinforcement Learning. https://openreview.net/forum?id=OhjGzRE5N6o

  20. [28]

    Minji Lee, Luiz Felipe Vecchietti, Hyunkyu Jung, Hyun Joo Ro, Meeyoung Cha, and Ho Min Kim. 2024. Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space. In Forty-first International Conference on Machine Learning. https://openreview.net...

  21. [29]

    Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives. 2023. Evolutionary-scale prediction of atomic-l...

  22. [30]

    Liu and Jorge Nocedal

    Dong C. Liu and Jorge Nocedal. 1989. On the limited memory BFGS method for large scale optimization. Mathematical Programming 45, 1 (01 Aug 1989), 503–528. https://doi.org/10.1007/BF01589116

  23. [31]

    Satvik Mehul Mashkaria, Siddarth Krishnamoorthy, and Aditya Grover. 2023. Generative Pretraining for Black-Box Optimization. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Bruns...

  24. [32]

    John Maynard Smith. 1970. Natural Selection and the Concept of a Protein Space. Nature 225, 5232 (Feb. 1970), 563–564. https://doi.org/10.1038/225563a0

  25. [33]

    Tung Nguyen, Sudhanshu Agrawal, and Aditya Grover. 2023. ExPT: Syn- thetic Pretraining for Few-Shot Experimental Design. In Advances in Neu- ral Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associa...

  26. [34]

    Yuchi Qiu and Guo-Wei Wei. 2022. CLADE 2.0: Evolution-Driven Cluster Learning-Assisted Directed Evolution. Journal of Chemical Information and Modeling 62, 19 (Sept. 2022), 4629–4641. https://doi.org/10.1021/acs.jcim.2c01046

  27. [35]

    Usha Nandini Raghavan, Réka Albert, and Soundar Kumara. 2007. Near linear time algorithm to detect community structures in large-scale networks. Phys. Rev. E 76 (Sep 2007), 036106. Issue 3. https://doi.org/10.1103/PhysRevE.76.036106

  28. [36]

    Zhizhou Ren, Jiahan Li, Fan Ding, Yuan Zhou, Jianzhu Ma, and Jian Peng. 2022. Proximal Exploration for Model-guided Protein Sequence Design. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , Kamalika ...

  29. [37]

    Sarkisyan, Dmitry A

    Karen S. Sarkisyan, Dmitry A. Bolotin, Margarita V. Meer, Dinara R. Usman- ova, Alexander S. Mishin, George V. Sharonov, Dmitry N. Ivankov, Nina G. Bozhanova, Mikhail S. Baranov, Onuralp Soylemez, Natalya S. Bogatyreva, Pe- ter K. Vlasov, Evgeny S. Egorov, Maria D. Logacheva, ...

  30. [38]

    Sam Sinai, Richard Wang, Alexander Whatley, Stewart Slocum, Elina Locane, and Eric D. Kelsic. 2020. AdaLead: A simple and robust adaptive greedy search algorithm for sequence design. CoRR abs/2010.02141 (2020). arXiv:2010.02141 https://arxiv.org/abs/2010.02141

  31. [39]

    Samuel Stanton, Wesley Maddox, Nate Gruver, Phillip Maffettone, Emily Delaney, Peyton Greenside, and Andrew Gordon Wilson. 2022. Accelerating Bayesian Optimization for Biological Sequence Design with Denoising Autoencoders. In Proceedings of the 39th International Conference o...

  32. [40]

    Brandon Trabucco, Xinyang Geng, Aviral Kumar, and Sergey Levine. 2022. Design- Bench: Benchmarks for Data-Driven Offline Model-Based Optimization. In Pro- ceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , K...

  33. [41]

    Brandon Trabucco, Aviral Kumar, Xinyang Geng, and Sergey Levine. 2021. Con- servative Objective Models for Effective Offline Model-Based Optimization. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) ,...

  34. [42]

    Thanh V. T. Tran and Truong Son Hy. 2024. Protein De- sign by Directed Evolution Guided by Large Language Mod- els. bioRxiv (2024). https://doi.org/10.1101/2023.11.28.568945 arXiv:https://www.biorxiv.org/content/early/2024/05/02/2023.11.28.568945.full.pdf

  35. [43]

    Lane, and Huimin Zhao

    Yajie Wang, Pu Xue, Mingfeng Cao, Tianhao Yu, Stephan T. Lane, and Huimin Zhao. 2021. Directed Evolution: Methodologies and Applications. Chemical Reviews 121, 20 (2021), 12384–12444. https://doi.org/10.1021/acs.chemrev.1c00260 arXiv:https://doi.org/10.1021/acs.chemrev.1c00260...

  36. [44]

    Wilson, Riccardo Moriconi, Frank Hutter, and Marc Peter Deisen- roth

    James T. Wilson, Riccardo Moriconi, Frank Hutter, and Marc Peter Deisen- roth. 2017. The reparameterization trick for acquisition functions. arXiv:1712.00424 [stat.ML] https://arxiv.org/abs/1712.00424

  37. [45]

    Dengyong Zhou, Olivier Bousquet, Thomas Lal, Jason Weston, and Bernhard Schölkopf. 2003. Learning with Local and Global Consistency. In Advances in Neural Information Processing Systems , S. Thrun, L. Saul, and B. Schölkopf (Eds.), Vol. 16. MIT Press. https://proceedings.neuri...

  38. [48]

    Thus, we have E∥𝑧−𝑥∥≤( 1−𝛽)E∥𝜖∥+ E∥𝑥∥ <(1−𝛽)√ 𝑑+ √ 𝑑 = 2(1−𝛽) √ 𝑑

    has shown that the upper bound of their expectations’ norms is √ 𝑑. Thus, we have E∥𝑧−𝑥∥≤( 1−𝛽)E∥𝜖∥+ E∥𝑥∥ <(1−𝛽)√ 𝑑+ √ 𝑑 = 2(1−𝛽) √ 𝑑. Therefore, E[𝐷(𝑧, Conv(X))] < 2(1−𝛽) √ 𝑑, which completes the proof. □ B Evaluation Metrics We provide mathematical definitions for each metri...

  39. [49]

    package to tune the hyperparameters for each task. The range of hyperparameters is listed below: • Number of nodes𝑁 :{4000, 5000, 6000,..., 20000}; • Coefficient𝛼 :{0.1, 0.15, 0.2,..., 0.9}; • Number of propagation layers𝑚 :{1, 2, 3, 4}; • Number of neighbors𝑘 :{2, 3, 4,..., 8...

  40. [2012]

    Foundations of Computa- tional mathematics 12, 6 (2012), 805–849

    The convex geometry of linear inverse problems. Foundations of Computa- tional mathematics 12, 6 (2012), 805–849

  41. [2023]

    Machine Learning: Science and Technology 4, 2 (April 2023), 025014

    Plug and play directed evolution of proteins with gradient-based discrete MCMC. Machine Learning: Science and Technology 4, 2 (April 2023), 025014. https://doi.org/10.1088/2632-2153/accacd

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.