REVIEW 4 major objections 5 minor 33 references
BoTier: Multi-Objective Bayesian Optimization with Tiered Composite Objectives
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read BoTier is a threshold-gated composite objective that benchmarks show reaches all objectives in fewer experiments than existing multi-objective Bayesian optimization methods.
desk verdict BoTier is a useful, well-packaged threshold scalarization for hierarchical MOO, but the exact objective saturates and is discontinuous, so the paper's framing and continuity claim need fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the BoTier scalarization $\Xi = \sum_{i=1}^N \min(\psi_i,t_i) \prod_{j<i} H(\psi_j-t_j)$, a sum of capped objectives gated by threshold crossings. The product of step functions $H(\psi_j-t_j)$ ensures a lower-priority objective contributes only after all higher-priority objectives meet their thresholds; once a subordinate objective is below its threshold, $\min(\psi_i,t_i)$ makes that objective the active driver, and when all thresholds are met, $\Xi$ accumulates the threshold constants and the top objective dominates. To make the score optimizable, the paper approximates $H$ and $\min$ by smooth logistic and softmax-like functions with a user-set sharpness $k$, and evaluates the composite objective through Monte-Carlo integration over posterior samples from independent Gaussian-process surrogates. That smooth-composite construction is what carries the claim: it preserves the ranking of Chimera in the reported correlation checks while making the objective auto-differentiable and batch-evaluable.
What would settle it
Run the same BoTier Bayesian-optimization campaigns on the four emulated chemistry problems with the exact, discontinuous score and with the smooth approximation at several $k$ values, and compare the fraction of campaigns that satisfy all thresholds within the 50-evaluation budget; if the smooth version underperforms the exact version on any problem, the equivalence assumption fails.
Extended reading notes
Core claim
The central claim is that tiered preferences in multi-objective optimization can be captured by the threshold-gated sum $\Xi$, and that optimizing $\Xi$ as a composite objective is both practically usable and sample-efficient. Unlike Chimera, whose value for one point depends on the other points in the batch through running maxima over the search space, $\Xi$ gives each candidate an independent score once thresholds are fixed. Replacing the Heaviside factors and the $\min$ with smooth, $k$-parameterized analogs makes the score differentiable, so it can be differentiated through Monte-Carlo samples of Gaussian-process posteriors and optimized with standard expected-improvement acquisition. The paper's benchmark-supported conclusion is that, for the problems studied, BoTier never performed worse and often performed notably better than existing multi-objective methods, especially when used as a composite objective.
Load-bearing premise
The benchmarks assume that the smooth stand-ins for the step function and the min function give the same ranking of candidate experiments as the exact, discontinuous threshold score, for every user-set threshold; the paper tests this sensitivity only on the analytic test surfaces, not on the chemistry emulators.
Editorial extensions
If this is right
- If BoTier works as claimed, experiment planners can encode a known priority structure over outcomes and input costs without mapping the full Pareto front, saving experimental budget on uninteresting trade-off regions.
- The composite formulation means input-dependent objectives enter the score directly, so a surrogate never has to relearn known correlations between inputs and objectives.
- Because the score is auto-differentiable and built on Monte-Carlo posterior samples, it plugs into standard Bayesian optimization loops with expected-improvement acquisition.
- The benchmark evidence implies that, for threshold-based hierarchies, hierarchical scalarization can find all-satisfying conditions faster than non-hierarchical penalty methods and Pareto-oriented EHVI.
- Even when all objectives depend only on experiment outputs, composite use of BoTier matched or beat black-box scalarization in the studied cases.
Reading between the lines
- The paper fixes the smoothness parameter $k$ and checks sensitivity only on analytic surfaces; a natural extension is an adaptive or annealed $k$ that sharpens as the budget grows, which could make the smooth score track the exact threshold objective more faithfully near crossings.
- The observation that joint multi-output surrogates rarely helped suggests that modeling each objective independently may be sufficient whenever a scalarization separates objectives; testing this on higher-dimensional objective spaces would tell whether the gain is intrinsic to composite scalarization or to the surrogate split.
- The 'never performed worse' statement is tied to the paper's threshold choices; thresholds that place many Pareto-optimal points near the satisfaction boundary could compress the ranking differences, so the claim should be read as benchmark-specific.
- The same gating construction could in principle be reused for hierarchical constraints or multi-fidelity settings, though the paper does not develop those directions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BoTier, a hierarchical scalarization composite objective for multi-objective Bayesian optimization. The scalarization in Eq. (2) is designed to reflect a user-defined hierarchy over objectives, where each objective contributes only if all superordinate objectives satisfy their thresholds; the paper claims this preserves continuity and is consistent with the existing Chimera scalarization. The authors provide an auto-differentiable implementation using smooth approximations of the Heaviside and min functions, integrate it with BoTorch, and benchmark it on analytical surfaces and emulated chemistry problems (Suzuki-Miyaura coupling, benzylation, enzymatic alkoxylation, silver nanoparticle synthesis). The central claims are that BoTier is a robust and flexible composite objective that never performs worse than existing methods and often performs notably better, especially when used as a composite objective.
Significance. If the claims hold, BoTier would be a practically useful contribution to multi-objective Bayesian optimization, particularly for autonomous experimentation, because it extends hierarchical scalarization to the composite-objective setting and provides an open-source, auto-differentiable implementation. The paper's strengths include the release of a PyPI/GitHub package, reproducible benchmark scripts, systematic comparisons against Chimera, penalty-based scalarization, and EHVI, and explicit acknowledgment of the role of the smoothing parameter. The benchmark suite covers both analytical and real-world emulated problems, and the paper is generally clearly written. However, the exact formulation in Eq. (2) and the actually implemented smooth version differ in properties that are load-bearing for the paper's stated claims, and the benchmark success metric is defined in terms of the very thresholds BoTier targets. These issues require careful treatment before the claims of general superiority can be accepted.
major comments (4)
- [§2.2, Eq. (2)] The claim that Eq. (2) preserves continuity is incorrect for the exact formulation. At a superordinate threshold crossing, e.g., as ψ1 crosses t1 in the N=2 case, the Heaviside product ∏_{j<i} H(ψ_j - t_j) switches the subordinate term min(ψ2,t2) on or off, producing a jump of magnitude approximately t2 (or min(ψ2,t2)). Additionally, on the region where all ψ_i ≥ t_i, Eq. (2) equals Σ_i t_i exactly, so the exact scalarization has zero gradient and cannot distinguish between a point that barely satisfies all thresholds and one that far exceeds the primary objective. The implementation described in SI Section 1.1 uses smooth sigmoid and soft-min approximations (SI Eqs. S3 and S4), which remove both the discontinuity and the flat saturation, but this means the benchmarked objective is not Eq. (2) as stated. The manuscript should either redefine BoTier as the smooth objective from the outset, clearly state the exact-vs-smooth distinction, or provide a rigorous argument for why the smooth version faithfully represents the intended hierarchy without introducing qualitatively different behavior.
- [§3, Figs. 2-3; SI §S2.2, §S3] The success metric in the benchmarks is the number of experiments required to satisfy the first n objectives (Fig. 2 bottom rows; Figs. S9-S18; Figs. S20-S26). This metric is exactly what the flat region of Eq. (2) is designed to satisfy, so the paper's claim that BoTier 'never performed worse, but often performed notably better' is only established for reaching thresholds, not for continued maximization of objectives above their thresholds. Since the stated task in the introduction includes maximizing objectives such as yield, the benchmarks should also report post-threshold progress (e.g., the best yield achieved among all feasible points within the budget). Without such measurements, the claim of general superiority over EHVI and other baselines is too broad.
- [SI §S1.1, Figs. S13 and S18] The sensitivity analyses for the smoothing parameter k (Figs. S13 and S18) are performed only on the analytical surfaces, not on the chemistry emulators. The paper's central benchmarks on real-life problems (Section 3, Fig. 3) therefore rely on an unverified assumption that the smooth approximations in SI Eqs. S3-S4 preserve the ranking and optimizer of the exact threshold objective for those emulator surfaces. The authors should either provide sensitivity results on the emulated problems or supply a formal or quantitative argument (e.g., Lipschitz-type bounds relating k to objective perturbations) that the analytical-surface results transfer. As written, the conclusion that BoTier is robust across all investigated cases is not fully supported by the presented experiments.
- [SI §S4, Table S11] The consistency check between BoTier and Chimera reports Spearman rank correlations of only 0.38 and 0.58 for the BNH and BNH* surfaces, which the text attributes to 'numerical inconsistency' near thresholds. Since these are exactly the surfaces where the paper's own benchmarks show BoTier diverging from Chimera most strongly, the dismissal is not adequately justified. The manuscript should quantify how the low rank correlation affects the optimization trajectories, or at least acknowledge that the two scalarizations can rank points very differently in threshold-dense regions, rather than presenting the overall high correlations as evidence that BoTier is equivalent to Chimera.
minor comments (5)
- [Abstract / §1] There is a typo in the first sentence: 'satisifies' should be 'satisfies'.
- [SI §S1.1] SI Eq. S4 is written for the max function, but the scalarization in Eq. (2) uses min. The text says 'analogs for the Heaviside step function and the max function'; it should either derive the soft-min explicitly or clarify that min(x1,x2) = -max(-x1,-x2) is used.
- [§3] The sentence 'In every case, using BoTier as a composite objective further accelerated optimization, which we initially attributed to the the surrogate model not needing to re-discover known correlations' contains a duplicated 'the'.
- [Fig. 2 and Fig. 3] The figures report the 'best observed value of Ξ', but Ξ in the implementation is the smooth approximation. The captions should state which version of Ξ is plotted, since the exact and smooth versions may differ noticeably in regions near thresholds.
- [SI §S2.2] The description of benchmark runs says 'Each optimization campaign starts by randomly drawing a single seed experiment'. Clarify whether the seed is drawn uniformly at random or via a fixed random seed for reproducibility, since 50 independent runs are reported.
Circularity Check
No circular derivation: BoTier is an explicitly defined scalarization, benchmarked against external baselines on external surfaces; minor self-citations are contextual and not load-bearing.
full rationale
The central object, Eq. (2), is a definition, not a fitted or derived quantity: Ξ = Σ_i min(ψ_i,t_i) ∏_{j<i} H(ψ_j−t_j) contains only user thresholds t_i and the smoothing parameter k, neither of which is calibrated to the benchmark outcomes. The paper's headline claim—'never performed worse, but often performed notably better than existing MOO methods'—is supported by comparisons against Chimera, EHVI, penalty-based scalarization, and Sobol sampling on external analytical surfaces and chemistry emulators; no benchmark result is constructed from the BoTier formula itself. Self-citations (refs. 9, 10, 22, and to BoTorch) provide context or methodology and are not used to justify the scalarization's validity. The SI's Chimera-consistency check (Spearman rank correlation) is a property check, not a derivation, and its low BNH value (0.38) is disclosed. Validation choices—defining success as threshold satisfaction and measuring consistency on the same surfaces—bias interpretation but do not make the derivation equivalent to its inputs. Separately, the exact Eq. (2) is discontinuous at superordinate threshold crossings and flat once all thresholds are met, contradicting the §2.2 continuity claim; this is a correctness risk about the exact vs. smoothed objective, not a circularity, so it does not raise the circularity score.
Assumptions & free parameters
free parameters (3)
- satisfaction thresholds t_i =
Per-surface hand picks: e.g., Suzuki-Miyaura yield > 65, cost < 3500, temperature < 85; BNH y0 > -60, y1 > -11; etc.
- smoothing parameter k =
No default stated; sensitivity shown in Figs. S13 and S18
- objective normalization ranges =
User-provided [0,1] ranges per objective
assumptions (5)
- ad hoc to paper Smooth sigmoid and soft-min analogs preserve the ranking and optimizer of the exact threshold objective.
- domain assumption The objective hierarchy and satisfaction thresholds are known a priori and correctly specified by the user.
- domain assumption Trained emulators (Random Forest or kNN) faithfully represent the real experimental response surfaces.
- domain assumption Independent single-task Gaussian processes are adequate surrogates for each objective.
- standard math Monte Carlo integration of posterior samples gives an unbiased estimate of the expected composite utility.
Cite this review
Pith. "Pith review of BoTier: Multi-Objective Bayesian Optimization with Tiered Composite Objectives." pith.science (2026). https://pith.science/paper/5IED26B5
@misc{pith2026250115554,
author = {Pith},
title = {Pith review of: BoTier: Multi-Objective Bayesian Optimization with Tiered Composite Objectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/5IED26B5}},
note = {Machine review of arXiv:2501.15554}
}
read the original abstract
Scientific optimization problems are usually concerned with balancing multiple competing objectives, which come as preferences over both the outcomes of an experiment (e.g. maximize the reaction yield) and the corresponding input parameters (e.g. minimize the use of an expensive reagent). Typically, practical and economic considerations define a hierarchy over these objectives, which must be reflected in algorithms for sample-efficient experiment planning. Herein, we introduce BoTier, a composite objective that can flexibly represent a hierarchy of preferences over both experiment outcomes and input parameters. We provide systematic benchmarks on synthetic and real-life surfaces, demonstrating the robust applicability of BoTier across a number of use cases. Importantly, BoTier is implemented in an auto-differentiable fashion, enabling seamless integration with the BoTorch library, thereby facilitating adoption by the scientific community.
Figures
Reference graph
Works this paper leans on
-
[1]
J. C. Fromer and C. W. Coley, Patterns, 2023, 4, 100678
work page 2023
-
[2]
A. S. Vel, D. Cortés-Borda and F.-X. Felpin, React. Chem. Eng., 2024, 9, 2882–2891
work page 2024
-
[3]
B. J. Shields, J. Stevens, J. Li, M. Parasram, F. Damani, J. I. Martinez Alvardo, J. M. Janey, R. P. Adams and A. G. Doyle,Nature, 2021, 590, 89–96
work page 2021
-
[4]
N. H. Angello, V . Rathore, W. Beker, A. Wołos, E. R. Jira, R. Roszak, T. C. Wu, C. M. Schroeder, A. Aspuru-Guzik, B. A. Grzybowski and M. D. Burke, Science, 2022, 378, 399–405
work page 2022
-
[5]
B. Shi, T. Lookman and D. Xue, Materials Genome Engineering Advances , 2023, 1, e14
work page 2023
- [6]
-
[7]
P. I. Frazier, A Tutorial on Bayesian Optimization , 2018, https://arxiv.org/abs/1807.02811
arXiv 2018
-
[8]
Garnett, Bayesian Optimization, Cambridge University Press, 2023
R. Garnett, Bayesian Optimization, Cambridge University Press, 2023. 7
work page 2023
Show all 33 references
-
[9]
G. Tom, S. P. Schmid, S. G. Baird, Y . Cao, K. Darvish, H. Hao, S. Lo, S. Pablo-García, E. M. Rajaonson, M. Skreta, N. Yoshikawa, S. Corapi, G. D. Akkoc, F. Strieth-Kalthoff, M. Seifrid and A. Aspuru-Guzik, Chem. Rev., 2024, 124, 9633–9732
2024
-
[10]
Strieth-Kalthoff, H
F. Strieth-Kalthoff, H. Hao, V . Rathore, J. Derasp, T. Gaudin, N. H. Angello, M. Seifrid, E. Trushina, M. Guy, J. Liu, X. Tang, M. Mamada, W. Wang, T. Tsagaantsooj, C. Lavigne, R. Pollice, T. C. Wu, K. Hotta, L. Bodo, S. Li, M. Haddadnia, A. Wołos, R. Roszak, C. T. Ser, C. Bo...
2024
-
[11]
A. D. Clayton, J. A. Manson, C. J. Taylor, T. W. Chamberlain, B. A. Taylor, G. Clemens and R. A. Bourne,React. Chem. Eng., 2019, 4, 1545–1554
2019
-
[12]
J. A. G. Torres, S. H. Lau, P. Anchuri, J. M. Stevens, J. E. Tabora, J. Li, A. Borovika, R. P. Adams and A. G. Doyle,J. Am. Chem. Soc., 2022, 144, 19999–20007
2022
-
[13]
C. J. Taylor, A. Pomberger, K. C. Felton, R. Grainger, M. Barecka, T. W. Chamberlain, R. A. Bourne, C. N. Johnson and A. A. Lapkin, Chem. Rev., 2023, 123, 3089–3126
2023
-
[14]
Waltz, IEEE Transactions on Automatic Control, 1967, 12, 179–180
F. Waltz, IEEE Transactions on Automatic Control, 1967, 12, 179–180
1967
-
[15]
Stadler, Multicriteria Optimization in Engineering and in the Sciences , Springer Nature, 1988
W. Stadler, Multicriteria Optimization in Engineering and in the Sciences , Springer Nature, 1988
1988
-
[16]
Rentmeesters, W
M. Rentmeesters, W. Tsai and K.-J. Lin, Proceedings of the 2nd IEEE International Conference on Engineering of Complex Computer Systems, 1996, pp. 76–79
1996
-
[17]
Deb, in Multi-objective Optimisation Using Evolutionary Algorithms: An Introduction , ed
K. Deb, in Multi-objective Optimisation Using Evolutionary Algorithms: An Introduction , ed. L. Wang, A. H. C. Ng and K. Deb, Springer London, London, 2011, pp. 3–34
2011
-
[18]
Emmerich, K
M. Emmerich, K. Giannakoglou and B. Naujoks, IEEE Transactions on Evolutionary Computation, 2006, pp. 421–439
2006
-
[19]
Daulton, M
S. Daulton, M. Balandat and E. Bakshy, Proceedings of the 34th International Conference on Neural Information Processing Systems, 2020
2020
-
[20]
Chugh, 2020 IEEE Congress on Evolutionary Computation, 2020, pp
T. Chugh, 2020 IEEE Congress on Evolutionary Computation, 2020, pp. 1–8
2020
-
[21]
Klarner, T
L. Klarner, T. G. J. Rudner, G. M. Morris, C. Deane and Y . W. Teh, Proceedings of the 41th International Conference on Machine Learning, 2024
2024
-
[22]
Kristiadi, F
A. Kristiadi, F. Strieth-Kalthoff, M. Skreta, P. Poupart, A. Aspuru-Guzik and G. Pleiss, Proceedings of the 41th International Conference on Machine Learning, 2024
2024
-
[23]
Astudillo and P
R. Astudillo and P. Frazier, Proceedings of the 36th International Conference on Machine Learning, 2019, pp. 354–363
2019
-
[24]
F. Häse, M. Aldeghi, R. J. Hickman, L. M. Roch, M. Christensen, E. Liles, J. E. Hein and A. Aspuru-Guzik, Mach. Learn. Sci. Technol., 2021, 2, 035021
2021
-
[25]
Balandat, B
M. Balandat, B. Karrer, D. Jiang, S. Daulton, B. Letham, A. G. Wilson and E. Bakshy, Advances in Neural Information Processing Systems, 2020, pp. 21524–21538
2020
-
[26]
Daulton, X
S. Daulton, X. Wan, D. Eriksson, M. Balandat, M. A. Osborne and E. Bakshy, Advances in Neural Information Processing Systems, 2022, pp. 12760 – 12774. 8
2022
-
[27]
B. E. Walker, J. H. Bannock, A. M. Nightingale and J. C. deMello, React. Chem. Eng., 2017, 2, 785–798
2017
-
[28]
M. T. M. Emmerich, A. H. Deutz and J. W. Klinkenberg, 2011 IEEE Congress of Evolutionary Computation (CEC), 2011, pp. 2147–2154
2011
-
[29]
F. Häse, L. M. Roch and A. Aspuru-Guzik, Chem. Sci., 2018, 9, 7642–7655
2018
-
[30]
Mekki-Berrada, Z
F. Mekki-Berrada, Z. Ren, T. Huang, W. K. Wong, F. Zheng, J. Xie, I. P. S. Tian, S. Jayavelu, Z. Mahfoud, D. Bash, K. Hippal- gaonkar, S. Khan, T. Buonassisi, Q. Li and X. Wang,npj Comput. Mater ., 2021, 7, 55
2021
-
[31]
A. M. Schweidtmann, A. D. Clayton, N. Holmes, E. Badford, R. A. Bourne and A. A. Lapkin,Chem. Eng. J., 2018, 352, 277–282
2018
-
[32]
https:github.com/fsk_lab/botier
-
[33]
smoothness
F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V . Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot and E. Duchesnay,J. Mach. Learn. Res., 2011, 12, 2825–2830. 9 Supplementary Materials ...
2011
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.