Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

Neural Model-based Optimization with Right-Censored Observations

T0 review · 4 major / 4 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read The paper claims that training neural-network surrogates with a Tobit loss, which treats right-censored observations as lower bounds, improves model-based optimization and achieves new best results on tuning a SAT solver's runtime and a…

desk verdict Sensible engineering combination of Tobit loss and NN ensembles for censored BO, but the evidence doesn't cleanly isolate the Tobit contribution; still deserves a thorough peer review. read the letter →

arxiv 2009.13828 v1 pith:JVNRAWW2 submitted 2020-09-29 cs.AI cs.LG

classification cs.AIcs.LG
keywords right-censoreddatamodel-basedoptimizationneuralnetworksTobitlossBayesianalgorithmconfigurationThompsonsamplingruntimeprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Model-based optimizers often speed up search by cutting off poorly performing runs early, producing right-censored observations: you only know the true runtime was at least as large as the cutoff. The paper argues that a regression model should treat these as lower bounds, not throw them away, and proposes a Tobit loss for neural-network surrogates that does exactly that. On synthetic benchmarks the Tobit-trained networks achieve lower prediction error than networks that ignore or drop censored data, and they roughly match an iterative imputation baseline at a fraction of the training cost. On two real optimization tasks, tuning the step count of a local-search SAT solver and the time-to-accuracy of neural networks, the approach outperforms the previous best model-based configurator.

What carries the argument

The load-bearing object is the Tobit loss of Eq. (4), a negative log-likelihood that switches between the normal density for uncensored observations and the normal survival term $1-\Phi(Z_i)$ for censored ones. This lets the network learn that a capped run's true value is above the cutoff, rather than taking the cutoff as the truth. An ensemble of networks provides predictive uncertainty, and Thompson sampling selects the next configuration so that only one network is trained per optimization iteration.

What would settle it

A controlled simulation in which the cutoff is deliberately made to depend on the run's partial progress (informative censoring) should show biased mean predictions from a Tobit-trained network; if the Tobit model is correct, prediction error should not change systematically with the strength of the dependence between progress and cutoff. A testable comparison would pit the Tobit loss against a model that explicitly conditions on the stopping rule on data generated under adaptive capping.

Watch

Extended reading notes

Core claim

The central claim is that a neural network trained with the Tobit loss (Eq. 4) can use right-censored observations from early-stopped runs as lower-bound information instead of discarding them or imputing them iteratively. The loss combines the normal density $\varphi(Z_i)$ for uncensored observations with the survival probability $1-\Phi(Z_i)$ for censored ones, so a censored point contributes the probability that the true value lies above the observed cutoff. The paper reports that these Tobit-trained networks, paired with a deep ensemble for uncertainty and Thompson sampling for acquisition, improve predictive accuracy on synthetic and real runtime data and achieve the best median performance on all three SAT-tuned instances and the best average rank across eight time-to-accuracy benchmarks.

Load-bearing premise

The load-bearing premise is that the stopping decision carries no information about the true runtime beyond the configuration itself, so a censored run's cutoff can be treated as an independent censoring threshold; in racing with adaptive capping the cutoff is chosen using the incumbent and the run's partial progress, which would make the Tobit estimate biased.

Editorial extensions

If this is right

  • Neural surrogates with the Tobit loss can replace random-forest models that need iterative imputation, removing a tuning hyperparameter and cutting training time by roughly the number of imputation iterations.
  • Censored observations become a usable training signal, so optimizers can cap runs aggressively without losing predictive accuracy.
  • The rank correlation between predicted and true configuration performance improves as censored data accumulate, which should translate into faster convergence of the search.
  • The approach generalizes across objective types, from low-dimensional SAT solver step counts to seven-dimensional neural network hyperparameters.
  • Thompson sampling with a single network per iteration keeps the per-step overhead small enough for practical model-based optimization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's Gaussian-on-log-runtime assumption is a modeling convenience; the same Tobit construction could be tried with heavy-tailed parametric families, and synthetic tests would show whether the gains persist.
  • Because the loss ignores how the cutoff is generated, the approach is safest when censoring is by a fixed global cutoff; extending it to model the racing rule jointly is a natural next step.
  • The time-to-accuracy benchmark introduced here could serve as a standard testbed for surrogate models under censoring beyond this paper.
  • A practical variant might draw multiple Thompson samples from a small ensemble rather than one network per iteration, trading a little overhead for better exploration in censored regions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper addresses right-censored observations in model-based optimization, where evaluations are terminated early and only a lower bound on the true response is observed. The authors propose training neural network surrogates with a Tobit loss (Eq. 4), which combines the normal density for uncensored observations with the survival probability 1-Phi(Z) for censored ones, and integrate this into a SMAC-style optimizer using an ensemble of networks and Thompson sampling. The paper reports predictive-quality experiments on synthetic functions and on runtime data from SAT solving and neural-network time-to-accuracy, and optimization experiments on the same two benchmarks, claiming new state-of-the-art performance.

Significance. The paper targets an important practical problem: model-based optimization with early stopping, where censored observations are common and neural-network surrogates are increasingly attractive. The Tobit loss is a natural and computationally cheap alternative to iterative Schmee-Hahn imputation, and the optimization results are evaluated with repeated independent runs of the final configurations, which is a real strength. If the concerns below are resolved, the paper would be a useful contribution. At present, however, the evidence that the Tobit loss itself drives the reported improvements is weakened by a likely formula error in the loss, an unaddressed informative-censoring problem in the racing setting, and a circular ground truth in the predictive-quality evaluation. No code is provided, which limits reproducibility.

major comments (4)
  1. [Training a Neural Network on Censored Observations, Eq. (3)] Equation (3) writes Z_i = (c_i - mu_i) / sigma_i^2, while the text says the second output neuron models the variance sigma_i^2. If sigma_i^2 denotes the variance, the standardized residual should be (c_i - mu_i) / sigma_i, and the uncensored contribution in Eq. (4) should be (1 / sigma_i) * phi((c_i - mu_i) / sigma_i). As written, the loss is not a proper negative log-likelihood: for uncensored points it contains no term that penalizes large variance, so sigma can grow without bound. Please correct the formula and state clearly whether the network parameterizes the standard deviation or the variance.
  2. [Formal Problem Setting, Eq. (2) and Eq. (4)] The Tobit likelihood in Eq. (4) treats the censoring threshold kappa_i as fixed and the censoring event as ignorable. In the adaptive-capping procedure described in the same section, kappa_i is chosen during the race and depends on the incumbent's observed performance and possibly on the partial progress of the current run, so the event c(lambda_i) > kappa_i is not independent of the latent runtime given lambda_i. The paper gives no identifiability argument or experiment showing that the Tobit MLE remains unbiased under such informative censoring. The synthetic experiment in 'Studying the Impact of Censored Observations' uses a fixed global threshold and does not reproduce the adaptive racing mechanism; please add an experiment with racing-style adaptive capping and uncensored ground truth, or provide explicit conditions under which the censoring is ignorable.
  3. [Model-based Optimization, footnote 2 in 'Evaluating the quality of our NNs on runtime data'] The predictive-quality comparison on actual runtime data derives ground-truth values for timed-out configurations by maximizing the Tobit likelihood on the same log-runtime values. Comparing a Tobit-trained model against 'ignore', 'drop', and Schmee-Hahn baselines using these Tobit-imputed labels is circular and favors the Tobit loss by construction. Please report results against ground truth obtained from uncensored repeated evaluations, at least on a subset of configurations, or explicitly characterize the comparison as relative rather than as ground-truth quality.
  4. [Model-based Optimization, Table 2] The optimization experiments compare the full NN+Thompson-sampling+Tobit pipeline with RF+Schmee-Hahn and random search. Because both the model class and the acquisition mechanism change, these results do not isolate the contribution of the Tobit loss. To support the claim that the Tobit loss is the driver of the improvement, include an ablation, for example NN+Thompson sampling trained without the Tobit correction (treating censored observations as observed, or dropping them).
minor comments (4)
  1. [Studying the Impact of Censored Observations, Problem Setup] The censoring rule 'with a probability increasing from 0 to 1' is ambiguous; please specify the functional form of the probability and how the observed censored value is generated.
  2. [References] The reference 'Pinitilie 2006' appears to be a typographical error for 'Pintilie 2006'; please correct it.
  3. [Conclusion] The abstract and conclusion state 'new state-of-the-art performance', but the comparison set is limited to one RF-based SMAC variant and random search; please either broaden the comparison or soften the claim.
  4. [Implementation Details] The paper states that code will be made publicly available upon acceptance; for reproducibility, please provide an anonymous repository or a detailed implementation appendix for the reported experiments.

Circularity Check

1 steps flagged · score 3.0 of 10

Predictive-quality comparison in Figure 5 uses Tobit-MLE-imputed ground truth, making that experiment partially self-referential; the headline optimization claim is independently evaluated.

  1. fitted input called prediction [Evaluating the quality of our NNs on runtime data, footnote 2 (used for Figure 5 ground truth); Eq. (4) defines the Tobit loss.]
    "To obtain results in a feasible time we applied a global cutoff yielding still some globally censored values. To obtain a ground truth value for each configuration to compare our prediction against, instead of using the empirical mean biased due to the global cutoff, we use the mean of a normal distribution fitted via maximizing the Tobit likelihood on the log-values. We note that this only makes a difference for configurations where we observed uncensored and timed-out runs. We replaced all values higher than the globally set cutoff by the cutoff before computing any metrics."

    Eq. (4) trains the network by minimizing the negative Tobit log-likelihood for a normal model of log-runtimes, using phi(Z_i) for uncensored and 1-Phi(Z_i) for censored observations. The footnote defines the evaluation target for censored configurations as the mean of a normal distribution fitted by maximizing the same Tobit likelihood on the same log-values. Thus, for configurations with timed-out runs, the 'ground truth' is not an independent measurement but the per-configuration Tobit MLE, which is the very quantity the proposed loss is designed to estimate. RMSE computed against these labels is therefore partially self-referential: ignoring or dropping censored data (I, D) is penalized relative to T partly because the target itself is defined by T's model family.

full rationale

The paper's central optimization claims rest on Table 2, where final configurations are re-evaluated with 1,000 (Saps) and 100 (NN) independent runs; those results are not circular. The Tobit loss itself is a standard external construction (Tobin 1958), and no load-bearing uniqueness theorem or ansatz is imported from the authors' prior work; self-citations (e.g., SMAC, RF+S&H) serve as framework and baselines, not as justification of the central mechanism. The synthetic experiments use an explicit censoring process and evaluate against the known synthetic function values, so they are self-contained. The only circular step is the Figure 5 evaluation: for configurations with censored runs, the ground-truth value is set to the per-configuration Tobit MLE of a normal distribution on log-values, while the proposed T method is trained with exactly that Tobit negative log-likelihood. This makes the RMSE/CC comparison partially favor T by construction for those configurations. The footnote openly acknowledges the imputation and its scope ('only makes a difference for configurations where we observed uncensored and timed-out runs'). This is a real but contained circularity in a supporting experiment; it does not force the Table 2 optimization results. The absence of an ablation isolating Tobit from NN+Thompson sampling is an attribution gap, not circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central method rests on a parametric normal assumption for log-runtimes and on independent right-censoring. No new entities are introduced. The hand-chosen network and training hyperparameters affect performance but are not fitted to the benchmark outcomes.

free parameters (4)
  • Network architecture (hidden layers, units, activation) = 3 layers x 50 units, tanh
    Taken from Snoek et al. (2015); hand-chosen capacity that affects model quality and optimization results.
  • Training hyperparameters (SGD momentum, batch size, cyclic LR, weight decay, gradient clip) = batch 16, max LR 1e-2, weight decay 1e-4, clip 0.1
    Hand-chosen settings in Implementation Details; not tuned on benchmark outcomes but influence method performance.
  • Ensemble size M (synthetic and model-quality experiments) = 5
    Hand-chosen; affects predictive variance estimates. The final optimizer uses one network per iteration, so this is not central to the optimization claim.
  • Schmee-Hahn baseline iterations = 5
    Baseline hyperparameter for the comparison, not part of the proposed method.
assumptions (4)
  • domain assumption Log-transformed runtimes are normally distributed with input-dependent mean and variance.
    Used to construct the negative log-likelihood and Tobit loss in Eqs. (3) and (4); the paper justifies it only by noting heavy tails are mitigated by the log transform.
  • domain assumption Censoring is non-informative: conditional on the configuration, the cutoff time kappa_i is independent of the true runtime c(lambda_i).
    Eq. (4) is the standard independent right-censoring likelihood; the actual racing and adaptive-capping mechanism in Figure 2 generates kappa_i from the incumbent and partial run progress, which is informative censoring.
  • domain assumption The observed censored value c_i = min(kappa_i, c(lambda_i)) is a lower bound on the true cost.
    Formal problem setting, Eq. (2); this is true by construction for right-censored observations.
  • domain assumption Costs are positive and the optimization objective is the expected cost E[c(lambda)].
    Eq. (1); standard in model-based optimization, but no justification is given that the mean is the right summary for heavy-tailed runtimes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Neural Model-based Optimization with Right-Censored Observations." pith.science (2026). https://pith.science/paper/JVNRAWW2

@misc{pith2026200913828,
  author       = {Pith},
  title        = {Pith review of: Neural Model-based Optimization with Right-Censored Observations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JVNRAWW2}},
  note         = {Machine review of arXiv:2009.13828}
}
read the original abstract

In many fields of study, we only observe lower bounds on the true response value of some experiments. When fitting a regression model to predict the distribution of the outcomes, we cannot simply drop these right-censored observations, but need to properly model them. In this work, we focus on the concept of censored data in the light of model-based optimization where prematurely terminating evaluations (and thus generating right-censored data) is a key factor for efficiency, e.g., when searching for an algorithm configuration that minimizes runtime of the algorithm at hand. Neural networks (NNs) have been demonstrated to work well at the core of model-based optimization procedures and here we extend them to handle these censored observations. We propose (i)~a loss function based on the Tobit model to incorporate censored samples into training and (ii) use an ensemble of networks to model the posterior distribution. To nevertheless be efficient in terms of optimization-overhead, we propose to use Thompson sampling s.t. we only need to train a single NN in each iteration. Our experiments show that our trained regression models achieve a better predictive quality than several baselines and that our approach achieves new state-of-the-art performance for model-based optimization on two optimization problems: minimizing the solution time of a SAT solver and the time-to-accuracy of neural networks.

Figures

Figures reproduced from arXiv: 2009.13828 by the authors.

Figure 2
Figure 2. Model-based algorithm configuration using racing [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Comparing the predictive quality of NNs ignoring censoring information (a) using S&H (b) using the Tobit Loss (c) [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Mean and variance of RMSE across a 5-fold CV of an ensemble built by using 5 iterations of the S&H algorithm and using the Tobit loss. Left: Hartmann6 with aggressive censoring starting at the 20th percentile and 52% censored data. Right: Hartmann6 with mild censoring starting at the 80th percentile and 19.5% censored data. at the overall results, T achieved the best results, not sur￾prisingly, followed by S&H since… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Results for I, D, S&H and T on actual runtime data. The upper plots show RMSE and CC when training on increasing [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 2
Figure 2. Figure 2: Results for I, D, S&H and T. We show predicted [PITH_FULL_IMAGE:figures/full_fig_p010_2.png]
Figure 1
Figure 1. Figure 1: Results for I, D, S&H and T on actual runtime [PITH_FULL_IMAGE:figures/full_fig_p010_1.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learned Offline Query Planning via Bayesian Optimization

    cs.DB 2025-02 conditional novelty 6.0 of 10

    BayesQO combines variational autoencoders and Bayesian optimization with censored timeouts to discover faster join-order plans offline for repetitive analytic workloads.

Reference graph

Works this paper leans on

66 extracted references · 64 canonical work pages · cited by 1 Pith paper

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter doi edition editor eid howpublished institution isbn issn journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Ans \'o tegui, C.; Malitsky, Y.; Sellmann, M.; and Tierney, K. 2015. Model-Based Genetic Algorithms for Algorithm Configuration. In Proc. of IJCAI '15 , 733--739

  4. [4]

    Brochu, E.; Cora, V.; and de Freitas, N. 2010. A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning. arXiv:1012.2599v1 [cs.LG]

  5. [5]

    P.; Bischl, B.; and St \" u tzle, T

    C \' a ceres, L. P.; Bischl, B.; and St \" u tzle, T. 2017. Evaluating random forest models for irace. In Proc. of GECCO , 1146--1153. ACM

  6. [6]

    Candelieri, A.; Perego, R.; and Archetti, F. 2018. Bayesian optimization of pump operations in water distribution systems. JGO 71(1): 213--235

  7. [7]

    Chen, J.; Mak, S.; Roshan, V.; and Zhang, C. 2019. Adaptive design for Gaussian process regression under censoring. arXiv preprint arXiv:1910.05452

  8. [8]

    Chiarandini, M.; Fawcett, C.; and Hoos, H. 2008. A Modular Multiphase Heuristic Solver for Post Enrolment Course Timetabling. In Proc. of PATAT '08

Show all 66 references
  1. [9]

    Coleman, C.; Kang, D.; Narayanan, D.; Nardi, L.; Zhao, T.; Zhang, J.; Bailis, P.; Olukotun, K.; R \' e , C.; and Zaharia, M. 2019. Analysis of DAWNBench , a Time-to-Accuracy Machine Learning Performance Benchmark. Operating Systems Review 53(1): 14--25

  2. [10]

    Cox, D. 1972. Regression Models and Life-Tables. Journal of the Royal Statistical Society 34: 187--220

  3. [11]

    H.; Hutter, F.; and Leyton - Brown, K

    Eggensperger, K.; Lindauer, M.; Hoos, H. H.; Hutter, F.; and Leyton - Brown, K. 2018. Efficient benchmarking of algorithm configurators via model-based surrogates. Machine Learning 107(1): 15--41

  4. [12]

    Eggensperger, K.; Lindauer, M.; and Hutter, F. 2018. Neural Networks for Predicting Algorithm Runtime Distributions. In Proc. of IJCAI '18 , 1442--1448

  5. [13]

    Falkner, S.; Klein, A.; and Hutter, F. 2018. BOHB : R obust and E fficient H yperparameter O ptimization at S cale. In Proc. of ICML '18 , 1437--1446

  6. [14]

    Feurer, M.; Klein, A.; Eggensperger, K.; Springenberg, J.; Blum, M.; and Hutter, F. 2015. Efficient and Robust Automated Machine Learning. In Proc. of N eur IPS '15 , 2962--2970

  7. [15]

    Feurer, M.; van Rijn, J.; Kadra, A.; Gijsbers, P.; Mallik, N.; Ravi, S.; Müller, A.; Vanschoren, J.; and Hutter, F. 2019. OpenML-Python : an extensible Python API for OpenML . arXiv:1911.02490 [cs.LG]

  8. [16]

    Franke, J.; Köhler, G.; Awad, N.; and Hutter, F. 2019. Neural Architecture Evolution in Deep Reinforcement Learning for Continuous Control. In MetaLearn'19

  9. [17]

    Gagliolo, M. 2010. Online Dynamic Algorithm Portfolios. Ph.D. thesis, Università della Svizzera Italiana

  10. [18]

    Gagliolo, M.; and Schmidhuber, J. 2006. Impact of censored sampling on the performance of restart strategies. In Proc. of CP '06 , 167--181

  11. [19]

    Gebser, M.; Kaminski, R.; Kaufmann, B.; Schaub, T.; Schneider, M.; and Ziller, S. 2011. A Portfolio Solver for Answer Set Programming: Preliminary Report. In Proc. of LPNMR '11 , 352--357

  12. [20]

    Gomes, C.; and Selman, B. 1997. Problem structure in the presence of perturbations. In Proc. of AAAI '97 , 221--226

  13. [21]

    Haider, H.; Hoehn, B.; Davis, S.; and Greiner, R. 2020. Effective Ways to Build and Evaluate Individual Survival Distributions. JMLR 21(85): 1--63

  14. [22]

    Huang, G.; Pleiss, G.; Liu, Z.; Hopcroft, J.; and Weinberger, K. 2017. Snapshot Ensembles: Train 1, get M for free. In Proc. of ICLR '17

  15. [23]

    Hutter, F. 2017. Towards true end-to-end learning & optimization. Invited talk held at the European Conference on Machine Learning & Principles and Practices of Knowledge Discovery in Databases ( ECML/PKDD '17)

  16. [24]

    Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2010. Automated Configuration of Mixed Integer Programming Solvers. In Proc. of CPAIOR '10 , 186--202

  17. [25]

    Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2011 a . Bayesian Optimization With Censored Response Data. In NeurIPS workshop on Bayesian Optimization, Sequential Experimental Design, and Bandits (BayesOpt'11)

  18. [26]

    Hutter, F.; Hoos, H.; and Leyton-Brown, K. 2011 b . Sequential Model-Based Optimization for General Algorithm Configuration. In Proc. of LION '11 , 507--523

  19. [27]

    Hutter, F.; Hoos, H.; Leyton-Brown, K.; and Murphy, K. 2010. Time-Bounded Sequential Parameter Optimization. In Proc. of LION '10 , 281--298

  20. [28]

    Hutter, F.; Hoos, H.; Leyton-Brown, K.; and St \"u tzle, T. 2009. Param ILS : An Automatic Algorithm Configuration Framework. JAIR 36: 267--306

  21. [29]

    Hutter, F.; Lindauer, M.; Balint, A.; Bayless, S.; Hoos, H.; and Leyton-Brown, K. 2017. The Configurable SAT Solver Challenge ( CSSC ). AIJ 243: 1--25

  22. [30]

    Hutter, F.; Tompkins, D.; and Hoos, H. 2002. Scaling and Probabilistic Smoothing: Efficient Dynamic Local Search for SAT . In Proc. of CP '02 , 233--248

  23. [31]

    Ishwaran, H.; Kogalur, U.; Blackstone, E.; and Lauer, M. 2008. Random survival forests. The Annals of Applied Statistics 2(3): 841--860

  24. [32]

    Jones, D.; Schonlau, M.; and Welch, W. 1998. Efficient Global Optimization of Expensive Black Box Functions. JGO 13: 455--492

  25. [33]

    Kalbfleisch, J.; and Prentice, R. 2002. The statistical analysis of failure time data, volume 360. John Wiley & Sons

  26. [34]

    Kaplan, E.; and Meier, P. 1958. Nonparametric Estimation from Incomplete Observations. Journal of the American Statistical Association 53: 457--481

  27. [35]

    Katzman, L.; Shaham, U.; Cloninger, A.; Brates, J.; Jiang, T.; and Kluger, Y. 2018. DeepSurv : personalized treatment recommender system using a Cox proportional hazards deep neural network. BMC Medical Research Methodology 18

  28. [36]

    Klein, J.; and Moeschberger, M. 2006. Survival analysis: techniques for censored and truncated data. Springer-Verlag

  29. [37]

    Kleinbaum, D.; and Klein, M. 2012. Survival analysis. Springer-Verlag

  30. [38]

    Kleinberg, R.; Leyton-Brown, K.; and Lucier, B. 2017. Efficiency Through Procrastination: Approximately Optimal Algorithm Configuration with Runtime Guarantees. In Proc. of IJCAI '17 , 2023--2031

  31. [39]

    Kleinberg, R.; Leyton - Brown, K.; Lucier, B.; and Graham, D. 2019. Procrastinating with Confidence: Near-Optimal, Anytime, Adaptive Algorithm Configuration. In Proc. of N eur IPS '19 , 8881--8891

  32. [40]

    Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017. Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. In Proc. of N eur IPS '17 , 6402--6413

  33. [41]

    Lee, C.; Zame, W.; Yoon, J.; and van der Schaar , M. 2018. DeepHit: A Deep Learning Approach to Survival Analysis with Competing Risks. In Proc. of AAAI '18 , 2314--2321

  34. [42]

    Li, Y.; Vinzamuri, B.; and Reddy, C. 2016. Regularized Weighted Linear Regression for High-dimensional Censored Data. In Proc. of SDM'16 , 45--53

  35. [43]

    Lindauer, M.; Eggensperger, K.; Feurer, M.; Falkner, S.; Biedenkapp, A.; and Hutter, F. 2017. SMAC v3: Algorithm configuration in P ython. github.com/automl/SMAC3

  36. [44]

    Lu, X.; and Roy, B. V. 2017. Ensemble sampling. In Proc. of N eur IPS '17 , 3258--3266

  37. [45]

    S.; Wu, C.; Xu, L.; Young, C.; and Zaharia, M

    Mattson, P.; Cheng, C.; Diamos, G.; Coleman, C.; Micikevicius, P.; Patterson, D.; Tang, H.; Wei, G.; Bailis, P.; Bittorf, V.; Brooks, D.; Chen, D.; Dutta, D.; Gupta, U.; Hazelwood, K.; Hock, A.; Huang, X.; Kang, D.; Kanter, D.; Kumar, N.; Liao, J.; Narayanan, D.; Oguntebi, T.;...

  38. [46]

    Mockus, J. 1994. Application of B ayesian approach to numerical methods of global and stochastic optimization. Journal of Global Optimization 4(4): 347--365

  39. [47]

    Neal, R. 1996. Bayesian Learning for Neural Networks. Lecture Notes in Computer Science. Springer-Verlag

  40. [48]

    Nix, D.; and Weigend, A. 1994. Estimating the mean and variance of the target probability distribution. In Proc. of ICNN '94 , 55--60

  41. [49]

    Osband, I.; Blundell, C.; Pritzel, A.; and Roy, B. V. 2016. Deep exploration via bootstrapped DQN. In Proc. of N eur IPS '16 , 4026--4034

  42. [50]

    Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019. PyTorch : An...

  43. [51]

    Perrone, V.; Jenatton, R.; Seeger, M.; and Archambeau, C. 2018. Scalable hyperparameter transfer learning. In Proc. of N eur IPS '18 , 12751--12761

  44. [52]

    Pinitilie, M. 2006. Competing Risks: A Practical Perspective. John Wiley & Sons

  45. [53]

    Schilling, N.; Wistuba, M.; Drumond, L.; and Schmidt-Thieme, L. 2015. Hyperparameter optimization with factorized multilayer perceptrons. In Proc. of ECML / PKDD '15 , 87--103

  46. [54]

    Schmee, J.; and Hahn, G. 1979. A simple method for regression analysis with censored data. Technometrics 21: 417--432

  47. [55]

    Shahriari, B.; Swersky, K.; Wang, Z.; Adams, R.; and de Freitas, N. 2016. Taking the Human Out of the Loop: A Review of B ayesian Optimization. Proceedings of the IEEE 104(1): 148--175

  48. [56]

    Smith, L. 2018. A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay. arXiv:1803.09820 [cs.LG]

  49. [57]

    Snoek, J.; Larochelle, H.; and Adams, R. 2012. Practical B ayesian Optimization of Machine Learning Algorithms. In Proc. of N eur IPS '12 , 2960--2968

  50. [58]

    Snoek, J.; Rippel, O.; Swersky, K.; Kiros, R.; Satish, N.; Sundaram, N.; Patwary, M.; Prabhat; and Adams, R. 2015. Scalable B ayesian Optimization Using Deep Neural Networks. In Proc. of ICML '15 , 2171--2180

  51. [59]

    Springenberg, J.; Klein, A.; Falkner, S.; and Hutter, F. 2016. Bayesian Optimization with Robust B ayesian Neural Networks. In Proc. of N eur IPS '16

  52. [60]

    Thompson, W. 1933. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika 25(3/4): 285--294

  53. [61]

    Tobin, J. 1958. Estimation of Relationships for Limited Dependent Variables. Econometrica 26(1): 24--36

  54. [62]

    Vallati, M.; Fawcett, C.; Gerevini, A.; Hoos, H.; and Saetti, A. 2013. Automatic Generation of Efficient Domain-Optimized Planners from Generic Parametrized Planners. In Proc. of SOCS '13

  55. [63]

    Vanschoren, J.; van Rijn, J.; Bischl, B.; and Torgo, L. 2014. OpenML : Networked Science in Machine Learning. SIGKDD Explor. Newsl. 15(2): 49--60

  56. [64]

    Weisz, G.; György, A.; and Szepesvári, C. 2019 a . CapsAndRuns: An Improved Method for Approximately Optimal Algorithm Configuration. In Proc. of ICML '19 , 6707--6715

  57. [65]

    Weisz, G.; György, A.; and Szepesvári, C. 2019 b . LeapsAndBounds: A Method for Approximately Optimal Algorithm Configuration. In Proc. of ICML '18 , 5257--5265

  58. [66]

    White, C.; Neiswanger, W.; and Savani, Y. 2019. Neural Architecture Search via Bayesian Optimization with a Neural Network Prior. In MetaLearn'19

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.