REVIEW 4 major objections 3 minor 49 references
GROOT: Effective Design of Biological Sequences with Limited Experimental Data
T0 review · 4 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A latent graph smoother turns scarce labels into reliable protein optimizers.
desk verdict A sensible, incremental LSO method whose empirical gains on scarce-label protein tasks are real, but whose central pseudo-label-fidelity assumption is never directly tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the latent kNN graph with interpolated nodes and label propagation. GROOT repeatedly samples a training embedding $x$ uniformly, draws noise $\epsilon \sim \mathcal{N}(0, I_d)$, and creates a synthetic node $z = \beta x + (1-\beta)\epsilon$ until the graph has $N$ nodes; edges are $k$-nearest neighbors under Euclidean distance, and fitness labels are updated over $m$ layers by $Y' = \alpha D^{-1/2}AD^{-1/2}Y + (1-\alpha)Y$. The $z$-formula is the exploration mechanism, label propagation is the pseudo-labeling mechanism, and the bound $2(1-\beta)\sqrt{d}$ from Proposition 2 is the reliability guarantee that keeps synthetic nodes inside a 'reliable zone' near the training hull.
What would settle it
Measure the correlation between Euclidean distance in the latent space and absolute fitness difference on held-out pairs from a protein family; if the correlation is near zero, the kNN label-propagation premise fails. A direct test: train GROOT on a family with a known one-mutation fitness cliff, and check whether the optimizer proposes the low-fitness mutant because a synthetic node near it received a high smoothed label.
Extended reading notes
Core claim
GROOT claims that a surrogate trained on pseudo-labeled synthetic latent points, generated by $z = \beta x + (1-\beta)\epsilon$ with $x$ a training embedding and $\epsilon \sim \mathcal{N}(0, I_d)$, can extrapolate beyond the training set while remaining reliable. The paper proves that, under its two assumptions, the synthetic points fall outside the training convex hull with probability tending to 1 as $d$ grows, and that their expected distance to the hull is bounded by $2(1-\beta)\sqrt{d}$. It then shows empirically that the smoothed surrogate, fitted to these labels with a simple MLP, enables gradient ascent and L-BFGS to beat previous baselines in the extreme-label regime of AAV and GFP, and that the same smoothing also improves a ReLSO surrogate.
Load-bearing premise
The method assumes that Euclidean distance in the ESM-2/VAE latent space tracks fitness similarity, so kNN label propagation over synthetic noisy nodes yields meaningful pseudo-labels.
Editorial extensions
If this is right
- With GROOT's smoothing, a two-layer MLP surrogate plus gradient ascent or L-BFGS outperforms AdaLead, CbAS, Bayesian optimization, GFN-AL, PEX, GGS, and ReLSO on the AAV harder1–harder3 and GFP harder1–harder3 benchmarks.
- Smoothing cuts surrogate error on train and holdout sets alike; on AAV harder1 the train MAE drops from 4.94 to 1.02 and the holdout MAE from 8.93 to 5.78.
- When the labeled subset is randomly sampled, 20% of the harder3 data, under 100 sequences, already matches the best point of the full harder3 set; in the lowest-fitness subsampling, roughly 50%, about 200 sequences, is enough.
- Adding the same smoothing to ReLSO raises its fitness by 60%, 64.7%, and 22.7% on AAV harder1, harder2, and harder3, indicating the mechanism transfers across surrogate architectures.
Reading between the lines
- If latent Euclidean distance is a good fitness-similarity proxy, the smoothing step could be attached to any pretrained protein language model encoder, not just the VAE used here; the paper does not test that substitution.
- The theory makes $\beta$ the dial between exploration and reliability, so tuning $\beta$ per protein family could further improve extrapolation; the paper leaves $\beta$ fixed and mentions mutation-effect analysis as future work.
- Because the encoder and surrogate are treated as interchangeable, the same graph-smoothing recipe should apply to RNA and small-molecule design tasks beyond the three Design-Bench domains tested.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses latent space optimization (LSO) for biological sequence design under extreme label scarcity. GROOT builds a kNN graph over ESM-2/VAE latent embeddings of the labeled training sequences, augments the graph with synthetic latent nodes formed by interpolating training latents with Gaussian noise, assigns these synthetic nodes pseudo-labels via one round of label propagation, trains a shallow MLP surrogate on the enlarged node set, and then runs gradient-based optimization (gradient ascent or L-BFGS) in the latent space. The authors report results on AAV and GFP fitness optimization at several difficulty levels and on three Design-Bench tasks, and they provide two propositions bounding the probability that synthetic nodes fall outside the training convex hull and the expected distance of such nodes to that hull. The paper claims state-of-the-art results across all difficulties on the two protein benchmarks and a 6x/1.3x fitness improvement over the training set in extreme low-data settings.
Significance. If the pseudo-labels assigned to the synthetic nodes are faithful, GROOT is a simple and computationally attractive data-augmentation recipe for the limited-label regime, and the paper's ablations (Tables 4 and 5) show that the smoothing step consistently reduces surrogate MAE and improves optimized fitness. The method is domain-agnostic in principle, and the authors evaluate on multiple benchmarks and release code, which are concrete strengths. However, the central mechanism -- label propagation over synthetic latent nodes producing informative fitness labels -- is not directly validated, and there is a mismatch between the implemented interpolation formula and the theoretically analyzed one. The SOTA claim is also stated more strongly than Table 3 supports unless non-degenerate diversity is explicitly made part of the criterion. These issues are fixable but require additional evidence or a substantial rewriting of the claims.
major comments (4)
- [Algorithm 1, line 4; Propositions 1-2] Algorithm 1, line 4 defines z as −β*x + (1−β)*ε, while Propositions 1 and 2 and the surrounding text (Section 3.3, "interpolating the learned latent x with random noise") define z = β*x + (1−β)*ε. The sign is geometrically material: as β→1, the implemented formula sends z toward −x, which is not an interpolation near a training node and can lie far outside the training hull. Proposition 2's bound on D(z, Conv(X)) therefore does not apply to the nodes that the algorithm actually generates. Either the algorithm must be changed to match the theory or the theory must be re-derived for the implemented formula; as written, the "reliable zone" justification does not cover the method being evaluated.
- [Algorithm 2 / Eq. (4); Table 4] The load-bearing assumption that label propagation assigns meaningful fitness values to synthetic nodes is never tested. With N=20,000 graph nodes and |D|≤1,157, roughly 94-98% of the surrogate training set is synthetic; Algorithm 2 initializes those nodes to 0 and, with the Section 4.1 settings m=1 and α=0.2, updates them once via Eq. (4). Their pseudo-labels are small weighted averages heavily diluted by zero-valued synthetic neighbors. Table 4 shows that smoothing reduces the trained surrogate's MAE on training and holdout sets, but that measures smoothness of the regression function, not whether the pseudo-labels recover oracle fitness at synthetic nodes. The authors should decode a random sample of synthetic latent nodes and compare their pseudo-labels to oracle scores, or perform a held-out label-removal experiment, to establish the correlation on which the entire method depends.
- [Table 3; Section 1.1] Section 1.1 claims GROOT "achieves state-of-the-art results across all difficulties" on the AAV and GFP benchmarks, but Table 3 reports ReLSO fitness 0.94 on GFP harder1, harder2, and harder3, while GROOT obtains 0.88, 0.87, and 0.62, respectively. The footnote excludes ReLSO because its generated population has collapsed to a single sequence (zero diversity), which is a legitimate quality concern. However, the paper does not define the SOTA criterion that would exclude a high-fitness but degenerate population. If the claim is "best non-degenerate designs," that should be stated explicitly; on the raw fitness metric the unqualified SOTA assertion is not supported.
- [Section 3.4; Assumption 2] Propositions 1 and 2 rely on Assumption 2, that latent vectors are i.i.d. N(0,I_d). The VAE loss in Eq. (2) only encourages this via a KL term with weight η; it does not guarantee the assumption for finite data or for the ESM-2-initialized encoder used here. Since these propositions are presented as the theoretical justification for label propagation on synthetic nodes, the authors should provide at least a diagnostic of Gaussianity (or a sensitivity analysis varying η and the KL weight) for the actual latent embeddings used in the main experiments. The current statement that the assumption is "achieved" by the KL loss is too strong without such evidence.
minor comments (3)
- [Section 4.1 vs. Table 8] Section 4.1 states that the hyperparameters "are not finely tuned" and sets the number of propagation layers to N_layers = 1, while Table 8 reports per-task Optuna tuning with, for example, m = 4 for the GFP tasks. The text should clarify which hyperparameter settings produced the numbers in Table 3 and whether the main results use the Section 4.1 defaults or the tuned values from Table 8.
- [Figure 2] Figure 2 plots distance to the set of training nodes, whereas Proposition 2 concerns distance to the convex hull. Since the convex hull contains the training set, the plotted quantity is a conservative check, but the caption and text should state this relationship explicitly to avoid appearing to validate the proposition with a different quantity.
- [Algorithm 1, line 4] Apart from the sign issue, the notation "V← −V∪{z}" is unorthodox and should be written as "V ← V ∪ {z}" to avoid confusing a set update with a sign operation.
Circularity Check
No significant circularity: GROOT's pseudo-label smoothing is a data-augmentation step evaluated against external oracles, and its theoretical propositions are independent Gaussian-geometry statements.
full rationale
I walked the derivation chain: latent embeddings -> synthetic nodes z = beta*x + (1-beta)*eps -> kNN graph -> label propagation -> surrogate training -> gradient-based MBO. The pseudo-labels are generated from the training labels by label propagation, so they are augmented training targets, not first-principles predictions; the reported fitness gains are measured by external oracles (GFP/AAV oracles from Kirjner et al.; exact Design-Bench oracles), so the central empirical claim does not reduce to the method's own inputs. Propositions 1-2 are mathematical statements about N(0,I_d) vectors and convex hulls under stated Assumptions 1-2; they are not fitted to benchmark outcomes, and Section 4.3 independently checks the 100% outside-hull claim and the distance upper bound. The only self-citation ([42], Tran & Hy 2024) appears in a general list of directed-evolution methods and is not load-bearing. The main text's 'not finely tuned' statement conflicts with Appendix E's Optuna tuning per task, and the paper never directly validates that Euclidean latent distance tracks fitness; these are empirical overfitting/validation caveats, not demonstrated equivalences. The Limitations section likewise flags hyperparameter sensitivity, which I weigh as a practical limitation rather than circularity.
Assumptions & free parameters
free parameters (8)
- beta (interpolation coefficient) =
not reported; 0.5 used in Proposition validation (Section 4.3)
- alpha (label propagation weight) =
0.2-0.6 per task (Table 8)
- N (number of graph nodes) =
4,000-20,000 per task (Table 8)
- m (number of label propagation layers) =
1-6 per task (Table 8)
- k (number of nearest neighbors) =
4-8 per task (Table 8)
- gamma (weighted adjacency factor) =
1.0
- d (latent dimension) =
320 for proteins, 128 for Design-Bench
- eta (VAE KL weight) =
not reported
assumptions (4)
- domain assumption Assumption 1: N << exp(d/(2(C_beta^2+2)))
- domain assumption Assumption 2: latent vectors x are i.i.d. N(0,I_d)
- domain assumption Euclidean latent distance correlates with fitness
- ad hoc to paper Label propagation on the constructed kNN graph yields accurate pseudo-labels for synthetic points
Cite this review
Pith. "Pith review of GROOT: Effective Design of Biological Sequences with Limited Experimental Data." pith.science (2026). https://pith.science/paper/SOUWMRWM
@misc{pith2026241111265,
author = {Pith},
title = {Pith review of: GROOT: Effective Design of Biological Sequences with Limited Experimental Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/SOUWMRWM}},
note = {Machine review of arXiv:2411.11265}
}
read the original abstract
Latent space optimization (LSO) is a powerful method for designing discrete, high-dimensional biological sequences that maximize expensive black-box functions, such as wet lab experiments. This is accomplished by learning a latent space from available data and using a surrogate model to guide optimization algorithms toward optimal outputs. However, existing methods struggle when labeled data is limited, as training the surrogate model with few labeled data points can lead to subpar outputs, offering no advantage over the training data itself. We address this challenge by introducing GROOT, a Graph-based Latent Smoothing for Biological Sequence Optimization. In particular, GROOT generates pseudo-labels for neighbors sampled around the training latent embeddings. These pseudo-labels are then refined and smoothed by Label Propagation. Additionally, we theoretically and empirically justify our approach, demonstrate GROOT's ability to extrapolate to regions beyond the training set while maintaining reliability within an upper bound of their expected distances from the training regions. We evaluate GROOT on various biological sequence design tasks, including protein optimization (GFP and AAV) and three tasks with exact oracles from Design-Bench. The results demonstrate that GROOT equalizes and surpasses existing methods without requiring access to black-box oracles or vast amounts of labeled data, highlighting its practicality and effectiveness. We release our code at https://anonymous.4open.science/r/GROOT-D554
Figures
Reference graph
Works this paper leans on
-
[1]
Michael Ahn, Henry Zhu, Kristian Hartikainen, Hugo Ponte, Abhishek Gupta, Sergey Levine, and Vikash Kumar. 2020. ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots. In Proceedings of the Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 100) , Leslie Pack Kael- bling, Danica Kragic, and Komei Sugiura (Eds.). PMLR, 1300...
work page 2020
-
[2]
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19). As- sociation for Computing Machinery, New York, NY, USA, 2623–2631. https...
arXiv 2019
-
[3]
Frances H. Arnold. 1996. Directed evolution: Creating biocatalysts for the future. Chemical Engineering Science 51, 23 (1996), 5091–5102. https://doi.org/10.1016/ S0009-2509(96)00288-6
work page 1996
-
[4]
Frances H. Arnold. 2018. Directed Evolution: Bringing New Chemistry to Life. Angewandte Chemie International Edition 57, 16 (2018), 4143–4148. https://doi.org/10.1002/anie.201708408 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/anie.201708408
-
[5]
Randall Balestriero, Jerome Pesenti, and Yann LeCun. 2021. Learning in high dimension always amounts to extrapolation. arXiv preprint arXiv:2110.09485 (2021)
arXiv 2021
-
[6]
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016. OpenAI Gym. arXiv:1606.01540 [cs.LG] https://arxiv.org/abs/1606.01540
arXiv 2016
-
[7]
David Brookes, Hahnbeom Park, and Jennifer Listgarten. 2019. Conditioning by adaptive sampling for robust design. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 773–782. https://proceedings.mlr.press/v97/brookes19a.html
work page 2019
-
[8]
Brookes, Amirali Aghazadeh, and Jennifer Listgarten
David H. Brookes, Amirali Aghazadeh, and Jennifer Listgarten. 2022. On the sparsity of fitness functions and implications for learning. Proceedings of the National Academy of Sciences 119, 1 (2022), e2109649118. https://doi.org/10.1073/ pnas.2109649118 arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.2109649118
Show all 49 references
-
[9]
David H Brookes, Amirali Aghazadeh, and Jennifer Listgarten. 2022. On the sparsity of fitness functions and implications for learning. Proceedings of the National Academy of Sciences 119, 1 (2022), e2109649118
2022
-
[10]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[11]
Bryant, Ali Bashir, Sam Sinai, Nina K
Drew H. Bryant, Ali Bashir, Sam Sinai, Nina K. Jain, Pierce J. Ogden, Patrick F. Riley, George M. Church, Lucy J. Colwell, and Eric D. Kelsic. 2021. Deep diversi- fication of an AAV capsid protein by machine learning. Nature Biotechnology 39, 6 (01 Jun 2021), 691–696. https://...
2021 doi
-
[12]
Egbert Castro, Abhinav Godavarthi, Julian Rubinfien, Kevin Givechian, Dhanan- jay Bhaskar, and Smita Krishnaswamy. 2022. Transformer-based protein genera- tion with regularized latent space optimization. Nature Machine Intelligence 4, 10 (01 Oct 2022), 840–851. https://doi.org...
2022 doi
-
[13]
Venkat Chandrasekaran, Benjamin Recht, Pablo A Parrilo, and Alan S Willsky
-
[14]
Can Chen, Yingxueff Zhang, Jie Fu, Xue (Steve) Liu, and Mark Coates. 2022. Bidirectional Learning for Offline Infinite-width Model-based Optimization. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds....
2022
-
[15]
Tianlai Chen, Pranay Vure, Rishab Pulugurta, and Pranam Chatterjee. 2023. AMP-Diffusion: Integrating Latent Diffusion with Protein Language Models for Antimicrobial Peptide Generation. In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop. https://openreview.net/forum?id...
2023
-
[16]
Patrick Emami, Aidan Perreault, Jeffrey Law, David Biagioni, and Peter St. John
-
[17]
Fowler and Stanley Fields
Douglas M. Fowler and Stanley Fields. 2014. Deep mutational scanning: a new style of protein science. Nature Methods 11, 8 (01 Aug 2014), 801–807. https: //doi.org/10.1038/nmeth.3027
2014 doi
-
[18]
Nathan C. Frey, Dan Berenberg, Karina Zadorozhny, Joseph Kleinhenz, Julien Lafrance-Vanasse, Isidro Hotzel, Yan Wu, Stephen Ra, Richard Bonneau, Kyunghyun Cho, Andreas Loukas, Vladimir Gligorijevic, and Saeed Saremi. 2024. Protein Discovery with Discrete Walk-Jump Sampling. In...
2024
-
[19]
Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D
Rafael Gómez-Bombarelli, Jennifer N. Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D. Hirzel, Ryan P. Adams, and Alán Aspuru- Guzik. 2018. Automatic Chemical Design Using a Data-Driven Conti...
2018 doi
-
[20]
Boyken, and David Baker
Po-Ssu Huang, Scott E. Boyken, and David Baker. 2016. The coming of age of de novo protein design. Nature 537, 7620 (01 Sep 2016), 320–327. https: //doi.org/10.1038/nature19946
2016 doi
-
[21]
Qian Huang, Horace He, Abhay Singh, Ser-Nam Lim, and Austin Benson. 2021. Combining Label Propagation and Simple Models out-performs Graph Neural Networks. In International Conference on Learning Representations . https:// openreview.net/forum?id=8E1-f3VhX1o
2021
-
[22]
Moksh Jain, Emmanuel Bengio, Alex Hernandez-Garcia, Jarrid Rector-Brooks, Bonaventure F. P. Dossou, Chanakya Ajit Ekbote, Jie Fu, Tianyu Zhang, Michael Kilgour, Dinghuai Zhang, Lena Simine, Payel Das, and Yoshua Bengio. 2022. Biological Sequence Design with GFlowNets. In Proce...
2022
-
[23]
Kauffman and Edward D
Stuart A. Kauffman and Edward D. Weinberger. 1989. The NK model of rugged fitness landscapes and its application to maturation of the immune response. Journal of Theoretical Biology 141, 2 (Nov. 1989), 211–245. https://doi.org/10. 1016/s0022-5193(89)80019-0
1989
-
[24]
Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. 2023. Segment Anything. arXiv:2304.02643 (2023)
2023 arXiv
-
[25]
Jaakkola, Regina Barzilay, and Ila R Fiete
Andrew Kirjner, Jason Yim, Raman Samusevich, Shahar Bracha, Tommi S. Jaakkola, Regina Barzilay, and Ila R Fiete. 2024. Improving protein optimization with smoothed fitness landscapes. In The Twelfth International Conference on Learning Representations. https://openreview.net/f...
2024
-
[26]
Aviral Kumar and Sergey Levine. 2020. Model Inversion Networks for Model- Based Optimization. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 5126–5137. https://proc...
2020
-
[27]
Minji Lee, Luiz Felipe Vecchietti, Hyunkyu Jung, Hyunjoo Ro, Meeyoung Cha, and Ho Min Kim. 2023. Protein Sequence Design in a Latent Space via Model-based Reinforcement Learning. https://openreview.net/forum?id=OhjGzRE5N6o
2023
-
[28]
Minji Lee, Luiz Felipe Vecchietti, Hyunkyu Jung, Hyun Joo Ro, Meeyoung Cha, and Ho Min Kim. 2024. Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space. In Forty-first International Conference on Machine Learning. https://openreview.net...
2024
-
[29]
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives. 2023. Evolutionary-scale prediction of atomic-l...
2023 doi
-
[30]
Liu and Jorge Nocedal
Dong C. Liu and Jorge Nocedal. 1989. On the limited memory BFGS method for large scale optimization. Mathematical Programming 45, 1 (01 Aug 1989), 503–528. https://doi.org/10.1007/BF01589116
1989 doi
-
[31]
Satvik Mehul Mashkaria, Siddarth Krishnamoorthy, and Aditya Grover. 2023. Generative Pretraining for Black-Box Optimization. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Bruns...
2023
-
[32]
John Maynard Smith. 1970. Natural Selection and the Concept of a Protein Space. Nature 225, 5232 (Feb. 1970), 563–564. https://doi.org/10.1038/225563a0
1970 doi
-
[33]
Tung Nguyen, Sudhanshu Agrawal, and Aditya Grover. 2023. ExPT: Syn- thetic Pretraining for Few-Shot Experimental Design. In Advances in Neu- ral Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associa...
2023
-
[34]
Yuchi Qiu and Guo-Wei Wei. 2022. CLADE 2.0: Evolution-Driven Cluster Learning-Assisted Directed Evolution. Journal of Chemical Information and Modeling 62, 19 (Sept. 2022), 4629–4641. https://doi.org/10.1021/acs.jcim.2c01046
2022 doi
-
[35]
Usha Nandini Raghavan, Réka Albert, and Soundar Kumara. 2007. Near linear time algorithm to detect community structures in large-scale networks. Phys. Rev. E 76 (Sep 2007), 036106. Issue 3. https://doi.org/10.1103/PhysRevE.76.036106
2007 doi
-
[36]
Zhizhou Ren, Jiahan Li, Fan Ding, Yuan Zhou, Jianzhu Ma, and Jian Peng. 2022. Proximal Exploration for Model-guided Protein Sequence Design. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , Kamalika ...
2022
-
[37]
Sarkisyan, Dmitry A
Karen S. Sarkisyan, Dmitry A. Bolotin, Margarita V. Meer, Dinara R. Usman- ova, Alexander S. Mishin, George V. Sharonov, Dmitry N. Ivankov, Nina G. Bozhanova, Mikhail S. Baranov, Onuralp Soylemez, Natalya S. Bogatyreva, Pe- ter K. Vlasov, Evgeny S. Egorov, Maria D. Logacheva, ...
2016
-
[38]
Sam Sinai, Richard Wang, Alexander Whatley, Stewart Slocum, Elina Locane, and Eric D. Kelsic. 2020. AdaLead: A simple and robust adaptive greedy search algorithm for sequence design. CoRR abs/2010.02141 (2020). arXiv:2010.02141 https://arxiv.org/abs/2010.02141
2020 arXiv
-
[39]
Samuel Stanton, Wesley Maddox, Nate Gruver, Phillip Maffettone, Emily Delaney, Peyton Greenside, and Andrew Gordon Wilson. 2022. Accelerating Bayesian Optimization for Biological Sequence Design with Denoising Autoencoders. In Proceedings of the 39th International Conference o...
2022
-
[40]
Brandon Trabucco, Xinyang Geng, Aviral Kumar, and Sergey Levine. 2022. Design- Bench: Benchmarks for Data-Driven Offline Model-Based Optimization. In Pro- ceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) , K...
2022
-
[41]
Brandon Trabucco, Aviral Kumar, Xinyang Geng, and Sergey Levine. 2021. Con- servative Objective Models for Effective Offline Model-Based Optimization. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) ,...
2021
-
[42]
Thanh V. T. Tran and Truong Son Hy. 2024. Protein De- sign by Directed Evolution Guided by Large Language Mod- els. bioRxiv (2024). https://doi.org/10.1101/2023.11.28.568945 arXiv:https://www.biorxiv.org/content/early/2024/05/02/2023.11.28.568945.full.pdf
2024 doi
-
[43]
Lane, and Huimin Zhao
Yajie Wang, Pu Xue, Mingfeng Cao, Tianhao Yu, Stephan T. Lane, and Huimin Zhao. 2021. Directed Evolution: Methodologies and Applications. Chemical Reviews 121, 20 (2021), 12384–12444. https://doi.org/10.1021/acs.chemrev.1c00260 arXiv:https://doi.org/10.1021/acs.chemrev.1c00260...
2021 doi
-
[44]
Wilson, Riccardo Moriconi, Frank Hutter, and Marc Peter Deisen- roth
James T. Wilson, Riccardo Moriconi, Frank Hutter, and Marc Peter Deisen- roth. 2017. The reparameterization trick for acquisition functions. arXiv:1712.00424 [stat.ML] https://arxiv.org/abs/1712.00424
2017 arXiv
-
[45]
Dengyong Zhou, Olivier Bousquet, Thomas Lal, Jason Weston, and Bernhard Schölkopf. 2003. Learning with Local and Global Consistency. In Advances in Neural Information Processing Systems , S. Thrun, L. Saul, and B. Schölkopf (Eds.), Vol. 16. MIT Press. https://proceedings.neuri...
2003
-
[48]
Thus, we have E∥𝑧−𝑥∥≤( 1−𝛽)E∥𝜖∥+ E∥𝑥∥ <(1−𝛽)√ 𝑑+ √ 𝑑 = 2(1−𝛽) √ 𝑑
has shown that the upper bound of their expectations’ norms is √ 𝑑. Thus, we have E∥𝑧−𝑥∥≤( 1−𝛽)E∥𝜖∥+ E∥𝑥∥ <(1−𝛽)√ 𝑑+ √ 𝑑 = 2(1−𝛽) √ 𝑑. Therefore, E[𝐷(𝑧, Conv(X))] < 2(1−𝛽) √ 𝑑, which completes the proof. □ B Evaluation Metrics We provide mathematical definitions for each metri...
2018
-
[49]
package to tune the hyperparameters for each task. The range of hyperparameters is listed below: • Number of nodes𝑁 :{4000, 5000, 6000,..., 20000}; • Coefficient𝛼 :{0.1, 0.15, 0.2,..., 0.9}; • Number of propagation layers𝑚 :{1, 2, 3, 4}; • Number of neighbors𝑘 :{2, 3, 4,..., 8...
2018
-
[2012]
Foundations of Computa- tional mathematics 12, 6 (2012), 805–849
The convex geometry of linear inverse problems. Foundations of Computa- tional mathematics 12, 6 (2012), 805–849
2012
-
[2023]
Machine Learning: Science and Technology 4, 2 (April 2023), 025014
Plug and play directed evolution of proteins with gradient-based discrete MCMC. Machine Learning: Science and Technology 4, 2 (April 2023), 025014. https://doi.org/10.1088/2632-2153/accacd
2023 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.