REVIEW 4 major objections 4 minor 1 cited by
Neural Model-based Optimization with Right-Censored Observations
T0 review · 4 major / 4 minor · reviewed 2026-08-27 · deepseek-v4-flash
Pith's one-line read The paper claims that training neural-network surrogates with a Tobit loss, which treats right-censored observations as lower bounds, improves model-based optimization and achieves new best results on tuning a SAT solver's runtime and a…
desk verdict Sensible engineering combination of Tobit loss and NN ensembles for censored BO, but the evidence doesn't cleanly isolate the Tobit contribution; still deserves a thorough peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Tobit loss of Eq. (4), a negative log-likelihood that switches between the normal density for uncensored observations and the normal survival term $1-\Phi(Z_i)$ for censored ones. This lets the network learn that a capped run's true value is above the cutoff, rather than taking the cutoff as the truth. An ensemble of networks provides predictive uncertainty, and Thompson sampling selects the next configuration so that only one network is trained per optimization iteration.
What would settle it
A controlled simulation in which the cutoff is deliberately made to depend on the run's partial progress (informative censoring) should show biased mean predictions from a Tobit-trained network; if the Tobit model is correct, prediction error should not change systematically with the strength of the dependence between progress and cutoff. A testable comparison would pit the Tobit loss against a model that explicitly conditions on the stopping rule on data generated under adaptive capping.
Extended reading notes
Core claim
The central claim is that a neural network trained with the Tobit loss (Eq. 4) can use right-censored observations from early-stopped runs as lower-bound information instead of discarding them or imputing them iteratively. The loss combines the normal density $\varphi(Z_i)$ for uncensored observations with the survival probability $1-\Phi(Z_i)$ for censored ones, so a censored point contributes the probability that the true value lies above the observed cutoff. The paper reports that these Tobit-trained networks, paired with a deep ensemble for uncertainty and Thompson sampling for acquisition, improve predictive accuracy on synthetic and real runtime data and achieve the best median performance on all three SAT-tuned instances and the best average rank across eight time-to-accuracy benchmarks.
Load-bearing premise
The load-bearing premise is that the stopping decision carries no information about the true runtime beyond the configuration itself, so a censored run's cutoff can be treated as an independent censoring threshold; in racing with adaptive capping the cutoff is chosen using the incumbent and the run's partial progress, which would make the Tobit estimate biased.
Editorial extensions
If this is right
- Neural surrogates with the Tobit loss can replace random-forest models that need iterative imputation, removing a tuning hyperparameter and cutting training time by roughly the number of imputation iterations.
- Censored observations become a usable training signal, so optimizers can cap runs aggressively without losing predictive accuracy.
- The rank correlation between predicted and true configuration performance improves as censored data accumulate, which should translate into faster convergence of the search.
- The approach generalizes across objective types, from low-dimensional SAT solver step counts to seven-dimensional neural network hyperparameters.
- Thompson sampling with a single network per iteration keeps the per-step overhead small enough for practical model-based optimization.
Reading between the lines
- The paper's Gaussian-on-log-runtime assumption is a modeling convenience; the same Tobit construction could be tried with heavy-tailed parametric families, and synthetic tests would show whether the gains persist.
- Because the loss ignores how the cutoff is generated, the approach is safest when censoring is by a fixed global cutoff; extending it to model the racing rule jointly is a natural next step.
- The time-to-accuracy benchmark introduced here could serve as a standard testbed for surrogate models under censoring beyond this paper.
- A practical variant might draw multiple Thompson samples from a small ensemble rather than one network per iteration, trading a little overhead for better exploration in censored regions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses right-censored observations in model-based optimization, where evaluations are terminated early and only a lower bound on the true response is observed. The authors propose training neural network surrogates with a Tobit loss (Eq. 4), which combines the normal density for uncensored observations with the survival probability 1-Phi(Z) for censored ones, and integrate this into a SMAC-style optimizer using an ensemble of networks and Thompson sampling. The paper reports predictive-quality experiments on synthetic functions and on runtime data from SAT solving and neural-network time-to-accuracy, and optimization experiments on the same two benchmarks, claiming new state-of-the-art performance.
Significance. The paper targets an important practical problem: model-based optimization with early stopping, where censored observations are common and neural-network surrogates are increasingly attractive. The Tobit loss is a natural and computationally cheap alternative to iterative Schmee-Hahn imputation, and the optimization results are evaluated with repeated independent runs of the final configurations, which is a real strength. If the concerns below are resolved, the paper would be a useful contribution. At present, however, the evidence that the Tobit loss itself drives the reported improvements is weakened by a likely formula error in the loss, an unaddressed informative-censoring problem in the racing setting, and a circular ground truth in the predictive-quality evaluation. No code is provided, which limits reproducibility.
major comments (4)
- [Training a Neural Network on Censored Observations, Eq. (3)] Equation (3) writes Z_i = (c_i - mu_i) / sigma_i^2, while the text says the second output neuron models the variance sigma_i^2. If sigma_i^2 denotes the variance, the standardized residual should be (c_i - mu_i) / sigma_i, and the uncensored contribution in Eq. (4) should be (1 / sigma_i) * phi((c_i - mu_i) / sigma_i). As written, the loss is not a proper negative log-likelihood: for uncensored points it contains no term that penalizes large variance, so sigma can grow without bound. Please correct the formula and state clearly whether the network parameterizes the standard deviation or the variance.
- [Formal Problem Setting, Eq. (2) and Eq. (4)] The Tobit likelihood in Eq. (4) treats the censoring threshold kappa_i as fixed and the censoring event as ignorable. In the adaptive-capping procedure described in the same section, kappa_i is chosen during the race and depends on the incumbent's observed performance and possibly on the partial progress of the current run, so the event c(lambda_i) > kappa_i is not independent of the latent runtime given lambda_i. The paper gives no identifiability argument or experiment showing that the Tobit MLE remains unbiased under such informative censoring. The synthetic experiment in 'Studying the Impact of Censored Observations' uses a fixed global threshold and does not reproduce the adaptive racing mechanism; please add an experiment with racing-style adaptive capping and uncensored ground truth, or provide explicit conditions under which the censoring is ignorable.
- [Model-based Optimization, footnote 2 in 'Evaluating the quality of our NNs on runtime data'] The predictive-quality comparison on actual runtime data derives ground-truth values for timed-out configurations by maximizing the Tobit likelihood on the same log-runtime values. Comparing a Tobit-trained model against 'ignore', 'drop', and Schmee-Hahn baselines using these Tobit-imputed labels is circular and favors the Tobit loss by construction. Please report results against ground truth obtained from uncensored repeated evaluations, at least on a subset of configurations, or explicitly characterize the comparison as relative rather than as ground-truth quality.
- [Model-based Optimization, Table 2] The optimization experiments compare the full NN+Thompson-sampling+Tobit pipeline with RF+Schmee-Hahn and random search. Because both the model class and the acquisition mechanism change, these results do not isolate the contribution of the Tobit loss. To support the claim that the Tobit loss is the driver of the improvement, include an ablation, for example NN+Thompson sampling trained without the Tobit correction (treating censored observations as observed, or dropping them).
minor comments (4)
- [Studying the Impact of Censored Observations, Problem Setup] The censoring rule 'with a probability increasing from 0 to 1' is ambiguous; please specify the functional form of the probability and how the observed censored value is generated.
- [References] The reference 'Pinitilie 2006' appears to be a typographical error for 'Pintilie 2006'; please correct it.
- [Conclusion] The abstract and conclusion state 'new state-of-the-art performance', but the comparison set is limited to one RF-based SMAC variant and random search; please either broaden the comparison or soften the claim.
- [Implementation Details] The paper states that code will be made publicly available upon acceptance; for reproducibility, please provide an anonymous repository or a detailed implementation appendix for the reported experiments.
Circularity Check
Predictive-quality comparison in Figure 5 uses Tobit-MLE-imputed ground truth, making that experiment partially self-referential; the headline optimization claim is independently evaluated.
-
fitted input called prediction
[Evaluating the quality of our NNs on runtime data, footnote 2 (used for Figure 5 ground truth); Eq. (4) defines the Tobit loss.]
"To obtain results in a feasible time we applied a global cutoff yielding still some globally censored values. To obtain a ground truth value for each configuration to compare our prediction against, instead of using the empirical mean biased due to the global cutoff, we use the mean of a normal distribution fitted via maximizing the Tobit likelihood on the log-values. We note that this only makes a difference for configurations where we observed uncensored and timed-out runs. We replaced all values higher than the globally set cutoff by the cutoff before computing any metrics."
Eq. (4) trains the network by minimizing the negative Tobit log-likelihood for a normal model of log-runtimes, using phi(Z_i) for uncensored and 1-Phi(Z_i) for censored observations. The footnote defines the evaluation target for censored configurations as the mean of a normal distribution fitted by maximizing the same Tobit likelihood on the same log-values. Thus, for configurations with timed-out runs, the 'ground truth' is not an independent measurement but the per-configuration Tobit MLE, which is the very quantity the proposed loss is designed to estimate. RMSE computed against these labels is therefore partially self-referential: ignoring or dropping censored data (I, D) is penalized relative to T partly because the target itself is defined by T's model family.
full rationale
The paper's central optimization claims rest on Table 2, where final configurations are re-evaluated with 1,000 (Saps) and 100 (NN) independent runs; those results are not circular. The Tobit loss itself is a standard external construction (Tobin 1958), and no load-bearing uniqueness theorem or ansatz is imported from the authors' prior work; self-citations (e.g., SMAC, RF+S&H) serve as framework and baselines, not as justification of the central mechanism. The synthetic experiments use an explicit censoring process and evaluate against the known synthetic function values, so they are self-contained. The only circular step is the Figure 5 evaluation: for configurations with censored runs, the ground-truth value is set to the per-configuration Tobit MLE of a normal distribution on log-values, while the proposed T method is trained with exactly that Tobit negative log-likelihood. This makes the RMSE/CC comparison partially favor T by construction for those configurations. The footnote openly acknowledges the imputation and its scope ('only makes a difference for configurations where we observed uncensored and timed-out runs'). This is a real but contained circularity in a supporting experiment; it does not force the Table 2 optimization results. The absence of an ablation isolating Tobit from NN+Thompson sampling is an attribution gap, not circularity.
Assumptions & free parameters
free parameters (4)
- Network architecture (hidden layers, units, activation) =
3 layers x 50 units, tanh
- Training hyperparameters (SGD momentum, batch size, cyclic LR, weight decay, gradient clip) =
batch 16, max LR 1e-2, weight decay 1e-4, clip 0.1
- Ensemble size M (synthetic and model-quality experiments) =
5
- Schmee-Hahn baseline iterations =
5
assumptions (4)
- domain assumption Log-transformed runtimes are normally distributed with input-dependent mean and variance.
- domain assumption Censoring is non-informative: conditional on the configuration, the cutoff time kappa_i is independent of the true runtime c(lambda_i).
- domain assumption The observed censored value c_i = min(kappa_i, c(lambda_i)) is a lower bound on the true cost.
- domain assumption Costs are positive and the optimization objective is the expected cost E[c(lambda)].
Cite this review
Pith. "Pith review of Neural Model-based Optimization with Right-Censored Observations." pith.science (2026). https://pith.science/paper/JVNRAWW2
@misc{pith2026200913828,
author = {Pith},
title = {Pith review of: Neural Model-based Optimization with Right-Censored Observations},
year = {2026},
howpublished = {\url{https://pith.science/paper/JVNRAWW2}},
note = {Machine review of arXiv:2009.13828}
}
read the original abstract
In many fields of study, we only observe lower bounds on the true response value of some experiments. When fitting a regression model to predict the distribution of the outcomes, we cannot simply drop these right-censored observations, but need to properly model them. In this work, we focus on the concept of censored data in the light of model-based optimization where prematurely terminating evaluations (and thus generating right-censored data) is a key factor for efficiency, e.g., when searching for an algorithm configuration that minimizes runtime of the algorithm at hand. Neural networks (NNs) have been demonstrated to work well at the core of model-based optimization procedures and here we extend them to handle these censored observations. We propose (i)~a loss function based on the Tobit model to incorporate censored samples into training and (ii) use an ensemble of networks to model the posterior distribution. To nevertheless be efficient in terms of optimization-overhead, we propose to use Thompson sampling s.t. we only need to train a single NN in each iteration. Our experiments show that our trained regression models achieve a better predictive quality than several baselines and that our approach achieves new state-of-the-art performance for model-based optimization on two optimization problems: minimizing the solution time of a SAT solver and the time-to-accuracy of neural networks.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Learned Offline Query Planning via Bayesian Optimization
BayesQO combines variational autoencoders and Bayesian optimization with censored timeouts to discover faster join-order plans offline for repetitive analytic workloads.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address author booktitle chapter doi edition editor eid howpublished institution isbn issn journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ans \'o tegui, C.; Malitsky, Y.; Sellmann, M.; and Tierney, K. 2015. Model-Based Genetic Algorithms for Algorithm Configuration. In Proc. of IJCAI '15 , 733--739
work page 2015
-
[4]
Brochu, E.; Cora, V.; and de Freitas, N. 2010. A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv:1012.2599v1 [cs.LG]
work page Pith review arXiv 2010
-
[5]
P.; Bischl, B.; and St \" u tzle, T
C \' a ceres, L. P.; Bischl, B.; and St \" u tzle, T. 2017. Evaluating random forest models for irace. In Proc. of GECCO , 1146--1153. ACM
work page 2017
-
[6]
Candelieri, A.; Perego, R.; and Archetti, F. 2018. Bayesian optimization of pump operations in water distribution systems. JGO 71(1): 213--235
work page 2018
-
[7]
Chen, J.; Mak, S.; Roshan, V.; and Zhang, C. 2019. Adaptive design for Gaussian process regression under censoring. arXiv preprint arXiv:1910.05452
work page Pith review arXiv 2019
-
[8]
Chiarandini, M.; Fawcett, C.; and Hoos, H. 2008. A Modular Multiphase Heuristic Solver for Post Enrolment Course Timetabling. In Proc. of PATAT '08
work page 2008
Show all 66 references
-
[9]
Coleman, C.; Kang, D.; Narayanan, D.; Nardi, L.; Zhao, T.; Zhang, J.; Bailis, P.; Olukotun, K.; R \' e , C.; and Zaharia, M. 2019. Analysis of DAWNBench , a Time-to-Accuracy Machine Learning Performance Benchmark. Operating Systems Review 53(1): 14--25
2019
-
[10]
Cox, D. 1972. Regression Models and Life-Tables. Journal of the Royal Statistical Society 34: 187--220
1972
-
[11]
H.; Hutter, F.; and Leyton - Brown, K
Eggensperger, K.; Lindauer, M.; Hoos, H. H.; Hutter, F.; and Leyton - Brown, K. 2018. Efficient benchmarking of algorithm configurators via model-based surrogates. Machine Learning 107(1): 15--41
2018
-
[12]
Eggensperger, K.; Lindauer, M.; and Hutter, F. 2018. Neural Networks for Predicting Algorithm Runtime Distributions. In Proc. of IJCAI '18 , 1442--1448
2018
-
[13]
Falkner, S.; Klein, A.; and Hutter, F. 2018. BOHB : R obust and E fficient H yperparameter O ptimization at S cale. In Proc. of ICML '18 , 1437--1446
2018
-
[14]
Feurer, M.; Klein, A.; Eggensperger, K.; Springenberg, J.; Blum, M.; and Hutter, F. 2015. Efficient and Robust Automated Machine Learning. In Proc. of N eur IPS '15 , 2962--2970
2015
-
[15]
Feurer, M.; van Rijn, J.; Kadra, A.; Gijsbers, P.; Mallik, N.; Ravi, S.; Müller, A.; Vanschoren, J.; and Hutter, F. 2019. OpenML-Python : an extensible Python API for OpenML . arXiv:1911.02490 [cs.LG]
2019 arXiv
-
[16]
Franke, J.; Köhler, G.; Awad, N.; and Hutter, F. 2019. Neural Architecture Evolution in Deep Reinforcement Learning for Continuous Control. In MetaLearn'19
2019
-
[17]
Gagliolo, M. 2010. Online Dynamic Algorithm Portfolios. Ph.D. thesis, Università della Svizzera Italiana
2010
-
[18]
Gagliolo, M.; and Schmidhuber, J. 2006. Impact of censored sampling on the performance of restart strategies. In Proc. of CP '06 , 167--181
2006
-
[19]
Gebser, M.; Kaminski, R.; Kaufmann, B.; Schaub, T.; Schneider, M.; and Ziller, S. 2011. A Portfolio Solver for Answer Set Programming: Preliminary Report. In Proc. of LPNMR '11 , 352--357
2011
-
[20]
Gomes, C.; and Selman, B. 1997. Problem structure in the presence of perturbations. In Proc. of AAAI '97 , 221--226
1997
-
[21]
Haider, H.; Hoehn, B.; Davis, S.; and Greiner, R. 2020. Effective Ways to Build and Evaluate Individual Survival Distributions. JMLR 21(85): 1--63
2020
-
[22]
Huang, G.; Pleiss, G.; Liu, Z.; Hopcroft, J.; and Weinberger, K. 2017. Snapshot Ensembles: Train 1, get M for free. In Proc. of ICLR '17
2017
-
[23]
Hutter, F. 2017. Towards true end-to-end learning & optimization. Invited talk held at the European Conference on Machine Learning & Principles and Practices of Knowledge Discovery in Databases ( ECML/PKDD '17)
2017
-
[24]
Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2010. Automated Configuration of Mixed Integer Programming Solvers. In Proc. of CPAIOR '10 , 186--202
2010
-
[25]
Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2011 a . Bayesian Optimization With Censored Response Data. In NeurIPS workshop on Bayesian Optimization, Sequential Experimental Design, and Bandits (BayesOpt'11)
2011
-
[26]
Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2011 b . Sequential Model-Based Optimization for General Algorithm Configuration. In Proc. of LION '11 , 507--523
2011
-
[27]
Hutter, F.; Hoos, H.; Leyton-Brown, K.; and Murphy, K. 2010. Time-Bounded Sequential Parameter Optimization. In Proc. of LION '10 , 281--298
2010
-
[28]
Hutter, F.; Hoos, H.; Leyton-Brown, K.; and St \"u tzle, T. 2009. Param ILS : An Automatic Algorithm Configuration Framework. JAIR 36: 267--306
2009
-
[29]
Hutter, F.; Lindauer, M.; Balint, A.; Bayless, S.; Hoos, H.; and Leyton-Brown, K. 2017. The Configurable SAT Solver Challenge ( CSSC ). AIJ 243: 1--25
2017
-
[30]
Hutter, F.; Tompkins, D.; and Hoos, H. 2002. Scaling and Probabilistic Smoothing: Efficient Dynamic Local Search for SAT . In Proc. of CP '02 , 233--248
2002
-
[31]
Ishwaran, H.; Kogalur, U.; Blackstone, E.; and Lauer, M. 2008. Random survival forests. The Annals of Applied Statistics 2(3): 841--860
2008
-
[32]
Jones, D.; Schonlau, M.; and Welch, W. 1998. Efficient Global Optimization of Expensive Black Box Functions. JGO 13: 455--492
1998
-
[33]
Kalbfleisch, J.; and Prentice, R. 2002. The statistical analysis of failure time data, volume 360. John Wiley & Sons
2002
-
[34]
Kaplan, E.; and Meier, P. 1958. Nonparametric Estimation from Incomplete Observations. Journal of the American Statistical Association 53: 457--481
1958
-
[35]
Katzman, L.; Shaham, U.; Cloninger, A.; Brates, J.; Jiang, T.; and Kluger, Y. 2018. DeepSurv : personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Medical Research Methodology 18
2018
-
[36]
Klein, J.; and Moeschberger, M. 2006. Survival analysis: techniques for censored and truncated data. Springer-Verlag
2006
-
[37]
Kleinbaum, D.; and Klein, M. 2012. Survival analysis. Springer-Verlag
2012
-
[38]
Kleinberg, R.; Leyton-Brown, K.; and Lucier, B. 2017. Efficiency Through Procrastination: Approximately Optimal Algorithm Configuration with Runtime Guarantees. In Proc. of IJCAI '17 , 2023--2031
2017
-
[39]
Kleinberg, R.; Leyton - Brown, K.; Lucier, B.; and Graham, D. 2019. Procrastinating with Confidence: Near-Optimal, Anytime, Adaptive Algorithm Configuration. In Proc. of N eur IPS '19 , 8881--8891
2019
-
[40]
Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. In Proc. of N eur IPS '17 , 6402--6413
2017
-
[41]
Lee, C.; Zame, W.; Yoon, J.; and van der Schaar , M. 2018. DeepHit: A Deep Learning Approach to Survival Analysis with Competing Risks. In Proc. of AAAI '18 , 2314--2321
2018
-
[42]
Li, Y.; Vinzamuri, B.; and Reddy, C. 2016. Regularized Weighted Linear Regression for High-dimensional Censored Data. In Proc. of SDM'16 , 45--53
2016
-
[43]
Lindauer, M.; Eggensperger, K.; Feurer, M.; Falkner, S.; Biedenkapp, A.; and Hutter, F. 2017. SMAC v3: Algorithm configuration in P ython. github.com/automl/SMAC3
2017
-
[44]
Lu, X.; and Roy, B. V. 2017. Ensemble sampling. In Proc. of N eur IPS '17 , 3258--3266
2017
-
[45]
S.; Wu, C.; Xu, L.; Young, C.; and Zaharia, M
Mattson, P.; Cheng, C.; Diamos, G.; Coleman, C.; Micikevicius, P.; Patterson, D.; Tang, H.; Wei, G.; Bailis, P.; Bittorf, V.; Brooks, D.; Chen, D.; Dutta, D.; Gupta, U.; Hazelwood, K.; Hock, A.; Huang, X.; Kang, D.; Kanter, D.; Kumar, N.; Liao, J.; Narayanan, D.; Oguntebi, T.;...
2017
-
[46]
Mockus, J. 1994. Application of B ayesian approach to numerical methods of global and stochastic optimization. Journal of Global Optimization 4(4): 347--365
1994
-
[47]
Neal, R. 1996. Bayesian Learning for Neural Networks. Lecture Notes in Computer Science. Springer-Verlag
1996
-
[48]
Nix, D.; and Weigend, A. 1994. Estimating the mean and variance of the target probability distribution. In Proc. of ICNN '94 , 55--60
1994
-
[49]
Osband, I.; Blundell, C.; Pritzel, A.; and Roy, B. V. 2016. Deep exploration via bootstrapped DQN. In Proc. of N eur IPS '16 , 4026--4034
2016
-
[50]
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019. PyTorch : An...
2019
-
[51]
Perrone, V.; Jenatton, R.; Seeger, M.; and Archambeau, C. 2018. Scalable hyperparameter transfer learning. In Proc. of N eur IPS '18 , 12751--12761
2018
-
[52]
Pinitilie, M. 2006. Competing Risks: A Practical Perspective. John Wiley & Sons
2006
-
[53]
Schilling, N.; Wistuba, M.; Drumond, L.; and Schmidt-Thieme, L. 2015. Hyperparameter optimization with factorized multilayer perceptrons. In Proc. of ECML / PKDD '15 , 87--103
2015
-
[54]
Schmee, J.; and Hahn, G. 1979. A simple method for regression analysis with censored data. Technometrics 21: 417--432
1979
-
[55]
Shahriari, B.; Swersky, K.; Wang, Z.; Adams, R.; and de Freitas, N. 2016. Taking the Human Out of the Loop: A Review of B ayesian Optimization. Proceedings of the IEEE 104(1): 148--175
2016
-
[56]
Smith, L. 2018. A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay. arXiv:1803.09820 [cs.LG]
2018 arXiv
-
[57]
Snoek, J.; Larochelle, H.; and Adams, R. 2012. Practical B ayesian Optimization of Machine Learning Algorithms. In Proc. of N eur IPS '12 , 2960--2968
2012
-
[58]
Snoek, J.; Rippel, O.; Swersky, K.; Kiros, R.; Satish, N.; Sundaram, N.; Patwary, M.; Prabhat; and Adams, R. 2015. Scalable B ayesian Optimization Using Deep Neural Networks. In Proc. of ICML '15 , 2171--2180
2015
-
[59]
Springenberg, J.; Klein, A.; Falkner, S.; and Hutter, F. 2016. Bayesian Optimization with Robust B ayesian Neural Networks. In Proc. of N eur IPS '16
2016
-
[60]
Thompson, W. 1933. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika 25(3/4): 285--294
1933
-
[61]
Tobin, J. 1958. Estimation of Relationships for Limited Dependent Variables. Econometrica 26(1): 24--36
1958
-
[62]
Vallati, M.; Fawcett, C.; Gerevini, A.; Hoos, H.; and Saetti, A. 2013. Automatic Generation of Efficient Domain-Optimized Planners from Generic Parametrized Planners. In Proc. of SOCS '13
2013
-
[63]
Vanschoren, J.; van Rijn, J.; Bischl, B.; and Torgo, L. 2014. OpenML : Networked Science in Machine Learning. SIGKDD Explor. Newsl. 15(2): 49--60
2014
-
[64]
Weisz, G.; György, A.; and Szepesvári, C. 2019 a . CapsAndRuns: An Improved Method for Approximately Optimal Algorithm Configuration. In Proc. of ICML '19 , 6707--6715
2019
-
[65]
Weisz, G.; György, A.; and Szepesvári, C. 2019 b . LeapsAndBounds: A Method for Approximately Optimal Algorithm Configuration. In Proc. of ICML '18 , 5257--5265
2019
-
[66]
White, C.; Neiswanger, W.; and Savani, Y. 2019. Neural Architecture Search via Bayesian Optimization with a Neural Network Prior. In MetaLearn'19
2019
Reviewed August 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.