REVIEW 4 major objections 6 minor 70 references
Neural Architecture Search by Estimation of Network Structure Distributions
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Irregular neural network architectures can be found with a probability matrix alone.
desk verdict Novel representation and honest limitations, but the EDA update's value over random sampling is unshown; worth refereeing, not citing yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The prototype matrix $P$ is the load-bearing object: a discrete probability distribution over layer types for every position in a growing feedforward network. The update rule (Eq. 2) re-estimates $P$ from the empirical layer choices of the best $K_s$ sampled networks, which is the same marginal-re-estimation step used by univariate estimation of distribution algorithms. Around this core, the paper adds three mechanisms: probability capping (clamping row entries to $[p_{\min}, p_{\max}]$), prototype inversion (replacing high probabilities with low ones when the mean row $L^2$-norm crosses a threshold, to escape premature convergence), and fixed shortcut patterns—residual or semi-dense—that are applied to every sampled network so deeper candidates train more reliably.
What would settle it
Take a sample of architectures from a late-stage prototype, train each for the 20-epoch brief regime and again for the 200-epoch full regime, and compute the rank correlation between the two accuracy orderings; if the correlation is near zero, the search's selection signal is dominated by training-speed artifacts rather than final model quality.
Extended reading notes
Core claim
The paper's central claim is that a single probability matrix—a prototype—can stand in for an entire population of network architectures and can be optimized to produce useful, non-regular CNNs. Under the assumption that each layer type is chosen independently, the prototype $P$ has one row per current layer and one column per operation in the layer library; sampling a network from $P$ gives a concrete feedforward architecture. After $K$ candidates are briefly trained and ranked, the top $K_s$ are used to set $P_{ij}=\frac{1}{|K_s|}\sum_{k=1}^{K_s} x^k_{ij}$, the empirical frequency of operation $j$ at layer $i$ among the selected models, and this is followed by appending newly initialized rows. The authors show that the resulting search discovers architectures without repeating operation sequences, such as a CIFAR-100 net dominated by large 7x7 and dilated 5x5 convolutions in later layers, and they report that these architectures are competitive with existing methods even though the search uses no weight sharing and only 20-epoch candidate training.
Load-bearing premise
The search assumes that the ranking of candidates after only 20 training epochs is accurate enough to guide re-estimation, and the paper concedes in Section IV-C that this condition does not strictly hold, with the baseline variant scoring best under brief training while ultimately performing worst.
Editorial extensions
If this is right
- The independence assumption plus a simple sampling loop is enough to find competitive CNNs, so the search space itself—not gradient-based architecture optimization or reinforcement learning—can carry much of the work.
- Irregular architectures that mix large kernels, dilated convolutions, and pooling at varying depths are reachable and can be competitive, so cell-based spaces may be leaving useful designs undiscovered.
- Because candidate networks are trained independently, the search parallelizes almost linearly with the number of GPUs.
- The algorithm's low-complexity bias means it naturally favors simpler, faster-to-train models, which is useful when a compact deployment target is more important than peak accuracy.
- Probability capping can stall the search, while inversion and shortcuts improve final accuracy, showing that convergence-control choices materially change the outcome.
Reading between the lines
- The brief-training rank distortion the authors document suggests a promising extension: replacing the fixed 20-epoch evaluation with adaptive budgets, such as spending more epochs on promising candidates, could improve final architectures without scaling cost linearly.
- Prototype inversion behaves like a tabu-style diversification mechanism; one could test whether inverting only the most certain rows (partial inversion) versus all rows is better on deeper searches, especially when combined with shortcut patterns.
- The same prototype representation could be applied to other structure-selection tasks, such as choosing operations in recurrent cells or transformer layers, where the independence assumption is even more approximate but may still provide a useful prior.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes ASED, an architecture search method that represents a CNN by a matrix of independent per-layer probabilities over ten layer types, updates the matrix with a UMDA-style re-estimation from the best Ks of K briefly trained candidates, and progressively adds layers. It reports experiments on USPS and CIFAR-100, comparing variants with probability capping, prototype inversion, and residual/semi-dense shortcut patterns, and it claims to discover irregular architectures competitive in accuracy and compute with cell-based NAS methods. The paper also includes an initialization baseline and a random-uniform architecture baseline.
Significance. If the empirical claims hold, ASED makes a useful contribution by demonstrating that a simple probability-matrix representation can search a non-cell, continuously growing architecture space, with interpretable state and easy parallelization. The authors are honest about limitations and provide source code. The reported best architectures exhibit genuinely non-repeating layer patterns that cannot be represented by identical cells, which is valuable. However, the strength of the causal claim that the EDA update drives search improvement is currently not established, and several comparisons lack the statistical and procedural controls needed to justify the word 'competitive.'
major comments (4)
- [Section IV-C, Fig. 3, Table 2] The paper explicitly states in Section IV-C that the assumption that brief-training rankings match full-training rankings 'does not strictly hold,' and the baseline variant has the highest brief-training validation accuracy but the weakest final accuracy. Since Eq. (2) re-estimates the prototype from the top-Ks briefly trained candidates, the central mechanism—that the prototype is tuned toward high-performance models—is not supported unless the selection signal is validated. I request a quantitative analysis of the rank correlation between 20-epoch and 200-epoch validation performance for a sample of architectures, or, alternatively, a reframing of the contribution as a search heuristic whose improvements come from the inversion/shortcut perturbations rather than from the EDA update.
- [Table 2] Every ASED variant is reported as a single best architecture from one stochastic search, with no repeated runs, confidence intervals, or significance tests. At 256 channels the ASED baseline (0.7483) is essentially matched by a single sample of 1000 networks from a uniform 16-layer prototype (0.7499), so the result does not demonstrate that iterative prototype updates improve over random sampling. The full-inversion variant's 0.7729 is promising, but without multiple seeds it is not possible to rule out seed luck. Please report repeated independent searches (mean and standard deviation, or at least best-of-k with error bars), and compare against random sampling under a matched computational budget.
- [Table 3] The comparison to published NAS methods in Table 3 is not made under a unified training setup; the text acknowledges that the results were obtained under non-matching environments and that some numbers are borrowed from other papers (e.g., PNAS/ENAS/DARTS/NAONet). Because final accuracy is highly sensitive to training schedule, regularization, and preprocessing, the claim that ASED is 'competitive both in accuracy and computational cost' is not established by this table. Please evaluate the final ASED architecture under the same training protocol as at least one strong competitor, or use a standardized benchmark such as NAS-Bench, or restrict the competitive claim to the internal baselines and non-cell-based methods.
- [Algorithm 1 and Section IV-A] Algorithm 1 returns the prototype P and says the final architecture is the one with highest probability, but the experiments report the 'best discovered architectures' and the 'best-performing network structures from each algorithm variant.' These are different selection rules: choosing the highest-probability architecture is a model-based output, while choosing the best validation-scored sampled network is best-of-search and inflates performance via selection bias. The paper must specify which rule was used and, if best-of-search, analyze the validation-gap/selection bias or use a hold-out selection procedure.
minor comments (6)
- [Equation (4)] The normalization formula is incomplete as written; define all symbols and specify how the row-wise scaling is applied when multiple entries are capped.
- [Section IV-A] 'PReLu' should be 'PReLU'; also 'We choose the follow' is a typo for 'We choose to follow.'
- [Figure 3] The caption references red dotted and green dashed lines but the text should state which colors correspond to max and median, and the y-axis label is missing units (accuracy).
- [Table 3] The row 'ASED (best)' should identify the variant (full inversion), the channel count (256), and the source of the 20-GPU-day search-cost estimate.
- [Section II] The related-work discussion would benefit from a direct comparison of ASED to PARSEC in the experiments or at least a clear statement of why PARSEC was not included in Table 3.
- [Section III-D] The shortcut patterns are fixed rules rather than learned structures; this should be stated more prominently in the abstract or contributions so that readers do not over-interpret 'non-linear connectivity' as optimized connectivity.
Circularity Check
No significant circularity: ASED is an empirical EDA-style search with external baselines; the admitted brief-training ranking distortion is a validity threat, not a circular derivation.
full rationale
This paper contains no mathematical derivation whose conclusion equals its premises; it is an empirical EDA-style search procedure. The prototype update in Eq. (2) is exactly the intended UMDA-style selection step, and the final architecture is read off the converged prototype; nothing is 'predicted' from a fitted quantity in the circular sense. The paper provides external baselines (best architecture from the initialization sample and best architecture from a uniformly random 16-layer prototype in Table 2), so the central comparison is falsifiable rather than constructed. The admitted failure of the brief-training ranking assumption in Section IV-C ('this condition does not strictly hold... validation results can be confusing for the search') is a threat to the validity of the selection signal and to the strength of the central claim, but it is an empirical assumption about training fidelity, not a definitional or self-citational circle. Self-citations (e.g., [12], [22], [45]-[47]) appear only as background and are not load-bearing for the ASED claim. No uniqueness theorem, ansatz smuggled through citation, or renamed known result was found. Although the random-uniform baseline nearly matches the baseline ASED variant at 256 channels, that is a weakness of the mechanism's demonstrated effect, not circularity.
Assumptions & free parameters
free parameters (5)
- Ninit =
5 (default), 2 (modified)
- K and Ks =
K=1000, Ks=100 (default); K=100, Ks=10 (modified)
- pmax =
0.9
- L2-norm inversion threshold =
0.65
- Shortcut parameter D =
2 or 3
assumptions (3)
- domain assumption Layer type choices are independent across positions (Section III-A).
- domain assumption Only layer types are optimized; filter sizes, strides, and channel counts are fixed (Section III-A).
- domain assumption Brief training (20 epochs) produces rankings that correlate with final performance (Section IV-A, IV-C).
Cite this review
Pith. "Pith review of Neural Architecture Search by Estimation of Network Structure Distributions." pith.science (2026). https://pith.science/paper/3YXVXCG3
@misc{pith2026190806886,
author = {Pith},
title = {Pith review of: Neural Architecture Search by Estimation of Network Structure Distributions},
year = {2026},
howpublished = {\url{https://pith.science/paper/3YXVXCG3}},
note = {Machine review of arXiv:1908.06886}
}
read the original abstract
The influence of deep learning is continuously expanding across different domains, and its new applications are ubiquitous. The question of neural network design thus increases in importance, as traditional empirical approaches are reaching their limits. Manual design of network architectures from scratch relies heavily on trial and error, while using existing pretrained models can introduce redundancies or vulnerabilities. Automated neural architecture design is able to overcome these problems, but the most successful algorithms operate on significantly constrained design spaces, assuming the target network to consist of identical repeating blocks. While such approach allows for faster search, it does so at the cost of expressivity. We instead propose an alternative probabilistic representation of a whole neural network structure under the assumption of independence between layer types. Our matrix of probabilities is equivalent to the population of models, but allows for discovery of structural irregularities, while being simple to interpret and analyze. We construct an architecture search algorithm, inspired by the estimation of distribution algorithms, to take advantage of this representation. The probability matrix is tuned towards generating high-performance models by repeatedly sampling the architectures and evaluating the corresponding networks, while gradually increasing the model depth. Our algorithm is shown to discover non-regular models which cannot be expressed via blocks, but are competitive both in accuracy and computational cost, while not utilizing complex dataflows or advanced training techniques, as well as remaining conceptually simple and highly extensible.
Figures
Reference graph
Works this paper leans on
-
[1]
Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 5 2015
work page 2015
-
[2]
Deep Residual Learning for Image Recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778
work page 2016
-
[3]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2261–2269
work page 2017
-
[4]
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 4, pp. 834–848, 2018
work page 2018
-
[5]
Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems 28, pp. 91–99, 2015
work page 2015
-
[6]
The Mythos of Model Interpretability,
Z. C. Lipton, “The Mythos of Model Interpretability,” in ICML 2016 Workshop on Human Interpretability in Machine Learning (WHI 2016) , 2016
work page 2016
-
[7]
Interpreting Deep Learning: The Machine Learning Rorschach Test?
A. S. Charles, “Interpreting Deep Learning: The Machine Learning Rorschach Test?”arXiv preprint arXiv:1806.00148, 2018
work page Pith review arXiv 2018
-
[8]
Methods for interpreting and understanding deep neural networks,
G. Montavon, W. Samek, and K.-R. Müller, “Methods for interpreting and understanding deep neural networks,” Digital Signal Processing, vol. 73, pp. 1–15, 2018
work page 2018
Show all 70 references
-
[9]
Deep Neural Networks Are Easily Fooled: High Confidence Predictions for Unrecognizable Images,
A. Nguyen, J. Yosinski, and J. Clune, “Deep Neural Networks Are Easily Fooled: High Confidence Predictions for Unrecognizable Images,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 427–436
2015
-
[10]
Identity mappings in deep residual networks,
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision (ECCV) , 2016, pp. 630–645
2016
-
[11]
A survey of transfer learning,
K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,”Journal of Big Data, vol. 3, no. 1, p. 9, 2016
2016
-
[12]
On the Layer Selection in Small-Scale Deep Networks,
A. Muravev, J. Raitoharju, and M. Gabbouj, “On the Layer Selection in Small-Scale Deep Networks,” in 7th European Workshop on Visual Information Processing (EUVIP), 2018
2018
-
[13]
A Deeper Look at Dataset Bias,
T. Tommasi, N. Patricia, B. Caputo, and T. Tuytelaars, “A Deeper Look at Dataset Bias,” in Domain Adaptation in Computer Vision Applications, G. Csurka, Ed. Cham: Springer International Publishing, 2017, pp. 37– 55
2017
-
[14]
Genetic algorithms and neural networks: optimizing connections and connectivity,
D. Whitley, T. Starkweather, and C. Bogart, “Genetic algorithms and neural networks: optimizing connections and connectivity,”Parallel Com- puting, vol. 14, no. 3, pp. 347–361, 1990
1990
-
[15]
Evolving artificial neural networks,
Xin Yao, “Evolving artificial neural networks,” Proceedings of the IEEE, vol. 87, no. 9, pp. 1423–1447, 1999
1999
-
[16]
Evolving Neural Network through Augmenting Topologies,
K. O. Stanley and R. Miikkulainen, “Evolving Neural Network through Augmenting Topologies,” Evolutionary Computation, vol. 10, no. 2, pp. 99–127, 2002
2002
-
[17]
A Hypercube-Based Encoding for Evolving Large-Scale Neural Networks,
K. O. Stanley, D. B. D’Ambrosio, and J. Gauci, “A Hypercube-Based Encoding for Evolving Large-Scale Neural Networks,” Artificial Life , vol. 15, no. 2, pp. 185–212, 2009. 14 VOLUME 9, 2016 Muravev et al.: Neural Architecture Search by Estimation of Network Structure Distributions
2009
-
[18]
Evolutionary artifi- cial neural networks by multi-dimensional particle swarm optimization,
S. Kiranyaz, T. Ince, A. Yildirim, and M. Gabbouj, “Evolutionary artifi- cial neural networks by multi-dimensional particle swarm optimization,” Neural Networks, vol. 22, no. 10, pp. 1448–1462, 2009
2009
-
[19]
Neuroevolution: from architec- tures to learning,
D. Floreano, P. Dürr, and C. Mattiussi, “Neuroevolution: from architec- tures to learning,”Evolutionary Intelligence, vol. 1, no. 1, pp. 47–62, 2008
2008
-
[20]
Genetic CNN,
L. Xie and A. Yuille, “Genetic CNN,” in IEEE International Conference on Computer Vision (ICCV), 2017, pp. 1388–1397
2017
-
[21]
Large-Scale Evolution of Image Classifiers,
E. Real et al., “Large-Scale Evolution of Image Classifiers,” in Interna- tional Conference on Machine Learning (ICML), 2017, pp. 2902–2911
2017
-
[22]
Finding Better Topologies for Deep Convolutional Neural Networks by Evolution,
H. Zhang, S. Kiranyaz, and M. Gabbouj, “Finding Better Topologies for Deep Convolutional Neural Networks by Evolution,” arXiv preprint arXiv:1809.03242, 2018
2018 arXiv
-
[23]
Regularized Evolution for Image Classifier Architecture Search,
E. Real, A. Aggarwal, Y . Huang, and Q. V . Le, “Regularized Evolution for Image Classifier Architecture Search,” Proceedings of the AAAI Confer- ence on Artificial Intelligence, vol. 33, pp. 4780–4789, 2019
2019
-
[24]
Evolving Deep Convolutional Neural Networks for Image Classification,
Y . Sun, B. Xue, M. Zhang, and G. G. Yen, “Evolving Deep Convolutional Neural Networks for Image Classification,” IEEE Transactions on Evolu- tionary Computation, vol. 24, pp. 394–407, 2019
2019
-
[25]
Neural Architecture Search with Reinforce- ment Learning,
B. Zoph and Q. V . Le, “Neural Architecture Search with Reinforce- ment Learning,” inInternational Conference on Learning Representations (ICLR), 2017
2017
-
[26]
Learning Transferable Architectures for Scalable Image Recognition,
B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning Transferable Architectures for Scalable Image Recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[27]
Efficient Neural Architecture Search via Parameter Sharing,
H. Pham, M. Y . Guan, B. Zoph, Q. V . Le, and J. Dean, “Efficient Neural Architecture Search via Parameter Sharing,” in Proceedings of the 35th International Conference on Machine Learning (PMLR), 2018, pp. 4095– 4104
2018
-
[28]
Progressive Neural Architecture Search,
C. Liu et al. , “Progressive Neural Architecture Search,” in European Conference on Computer Vision (ECCV). Cham: Springer International Publishing, 2018, pp. 19–35
2018
-
[29]
Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution,
T. Elsken, J. H. Metzen, and F. Hutter, “Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution,” in International Confer- ence on Learning Representations (ICLR), 2019
2019
-
[30]
DARTS: Differentiable Architecture Search,
H. Liu, K. Simonyan, and Y . Yang, “DARTS: Differentiable Architecture Search,” inInternational Conference on Learning Representations (ICLR), 2019
2019
-
[31]
Evaluat- ing the Search Phase of Neural Architecture Search,
C. Sciuto, K. Yu, M. Jaggi, C. Musat, and M. Salzmann, “Evaluat- ing the Search Phase of Neural Architecture Search,” arXiv preprint arXiv:1902.08142, 2019
1902 arXiv
-
[32]
Population-Based Incremental Learning: A Method for In- tegrating Genetic Search Based Function Optimization and Competitive Learning,
S. Baluja, “Population-Based Incremental Learning: A Method for In- tegrating Genetic Search Based Function Optimization and Competitive Learning,” Pittsburgh, PA, USA, 1994
1994
-
[33]
The Equation for Response to Selection and Its Use for Prediction,
H. Mühlenbein, “The Equation for Response to Selection and Its Use for Prediction,”Evolutionary Computation, vol. 5, no. 3, pp. 303–346, 1997
1997
-
[34]
Learning Represen- tations by Back-propagating Errors,
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning Represen- tations by Back-propagating Errors,” in Neurocomputing: Foundations of Research, J. A. Anderson and E. Rosenfeld, Eds. Cambridge, MA, USA: MIT Press, 1988, pp. 696–699
1988
-
[35]
Evolutionary computation: comments on the history and current state,
T. Back, U. Hammel, and H.-P. Schwefel, “Evolutionary computation: comments on the history and current state,” IEEE Transactions on Evo- lutionary Computation, vol. 1, no. 1, pp. 3–17, 1997
1997
-
[36]
Luke, Essentials of Metaheuristics, 2nd ed
S. Luke, Essentials of Metaheuristics, 2nd ed. Lulu, 2013
2013
-
[37]
Incremental Evolution of Complex General Behavior,
F. Gomez and R. Miikkulainen, “Incremental Evolution of Complex General Behavior,”Adaptive Behavior, vol. 5, no. 3-4, pp. 317–342, 1997
1997
-
[38]
A new evolutionary system for evolving artificial neural networks,
X. Yao and Y . Liu, “A new evolutionary system for evolving artificial neural networks,” IEEE Transactions on Neural Networks , vol. 8, no. 3, pp. 694–713, 5 1997
1997
-
[39]
Solving non-Markovian Control Tasks with Neuroevolution,
F. J. Gomez and R. Miikkulainen, “Solving non-Markovian Control Tasks with Neuroevolution,” in Proceedings of the 16th International Joint Conference on Artificial Intelligence - Volume 2, 1999, pp. 1356–1361
1999
-
[40]
Compositional pattern producing networks: A novel abstraction of development,
K. O. Stanley, “Compositional pattern producing networks: A novel abstraction of development,” Genetic Programming and Evolvable Ma- chines, vol. 8, no. 2, pp. 131–162, 2007
2007
-
[41]
An Enhanced Hypercube-Based Encoding for Evolving the Placement, Density, and Connectivity of Neurons,
S. Risi and K. O. Stanley, “An Enhanced Hypercube-Based Encoding for Evolving the Placement, Density, and Connectivity of Neurons,”Artificial Life, vol. 18, no. 4, pp. 331–363, 2012
2012
-
[42]
HyperNEAT: The First Five Years,
D. B. D’Ambrosio, J. Gauci, and K. O. Stanley, “HyperNEAT: The First Five Years,” in Growing Adaptive Machines . Springer, Berlin, Heidelberg, 2014, pp. 159–185
2014
-
[43]
Evolving Deep Neural Networks,
R. Miikkulainen et al., “Evolving Deep Neural Networks,” inArtificial In- telligence in the Age of Neural Networks and Brain Computing. Academic Press, 2019, pp. 293 – 312
2019
-
[44]
Completely Automated CNN Architecture Design Based on Blocks,
Y . Sun, B. Xue, M. Zhang, and G. G. Yen, “Completely Automated CNN Architecture Design Based on Blocks,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–13, 2019
2019
-
[45]
Progressive Opera- tional Perceptrons,
S. Kiranyaz, T. Ince, A. Iosifidis, and M. Gabbouj, “Progressive Opera- tional Perceptrons,”Neurocomputing, vol. 224, pp. 142–154, 2017
2017
-
[46]
Operational Neural Networks,
——, “Operational Neural Networks,” Neural Computing and Applica- tions (in print), 2020
2020
-
[47]
Heterogeneous Multilayer Generalized Operational Perceptron,
D. T. Tran, S. Kiranyaz, M. Gabbouj, and A. Iosifidis, “Heterogeneous Multilayer Generalized Operational Perceptron,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–15, 2019
2019
-
[48]
SMASH: One-Shot Model Architecture Search through HyperNetworks,
A. Brock, T. Lim, J. M. Ritchie, and N. Weston, “SMASH: One-Shot Model Architecture Search through HyperNetworks,” in Workshop on Meta-Learning (MetaLearn 2017) at NIPS, 2017
2017
-
[49]
Simple and Efficient Architecture Search for Convolutional Neural Networks,
T. Elsken, J.-H. Metzen, and F. Hutter, “Simple and Efficient Architecture Search for Convolutional Neural Networks,” in 6th International Confer- ence on Learning Representations (ICLR), 2018
2018
-
[50]
In- staNAS: Instance-aware Neural Architecture Search,
A.-C. Cheng, C. H. Lin, D.-C. Juan, W. Wei, and M. Sun, “In- staNAS: Instance-aware Neural Architecture Search,” arXiv preprint arXiv:1811.10201, 2018
2018 arXiv
-
[51]
Neural Architecture Search with Bayesian Optimisation and Optimal Transport,
K. Kandasamy, W. Neiswanger, J. Schneider, B. Poczos, and E. P. Xing, “Neural Architecture Search with Bayesian Optimisation and Optimal Transport,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 2016–2025
2018
-
[52]
Probabilistic Neural Architecture Search,
F. P. Casale, J. Gordon, and N. Fusi, “Probabilistic Neural Architecture Search,”arXiv preprint arXiv:1902.05116, 2019
1902 arXiv
-
[53]
A Survey of Optimization by Building and Using Probabilistic Models,
M. Pelikan, D. E. Goldberg, and F. G. Lobo, “A Survey of Optimization by Building and Using Probabilistic Models,” Computational Optimization and Applications, vol. 21, no. 1, pp. 5–20, 2002
2002
-
[54]
An introduction and survey of estimation of distribution algorithms,
M. Hauschild and M. Pelikan, “An introduction and survey of estimation of distribution algorithms,”Swarm and Evolutionary Computation, vol. 1, no. 3, pp. 111–128, 2011
2011
-
[55]
Comparing two K-category assignments by a K-category correlation coefficient,
J. Gorodkin, “Comparing two K-category assignments by a K-category correlation coefficient,” Computational Biology and Chemistry , vol. 28, no. 5-6, pp. 367–374, 12 2004
2004
-
[56]
On Stability of Fixed Points of Limit Models of Univariate Marginal Distribution Algorithm and Factorized Distribution Algorithm,
Q. Zhang, “On Stability of Fixed Points of Limit Models of Univariate Marginal Distribution Algorithm and Factorized Distribution Algorithm,” IEEE Transactions on Evolutionary Computation, vol. 8, no. 1, pp. 80–93, 2004
2004
-
[57]
EDAs Cannot Be Balanced and Stable,
T. Friedrich, T. Kötzing, and M. S. Krejca, “EDAs Cannot Be Balanced and Stable,” inProceedings of the Genetic and Evolutionary Computation Conference, 2016, pp. 1139–1146
2016
-
[58]
Tabu Search,
F. Glover and M. Laguna, “Tabu Search,” in Handbook of Combinatorial Optimization, D.-Z. Du and P. M. Pardalos, Eds. Boston, MA: Springer US, 1998, pp. 2093–2229
1998
-
[59]
A Database for Handwritten Text Recognition Research,
J. J. Hull, “A Database for Handwritten Text Recognition Research,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 1994
1994
-
[60]
Learning multiple layers of features from tiny images,
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009
2009
-
[61]
Delving deep into rectifiers: Surpass- ing human-level performance on imagenet classification,
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpass- ing human-level performance on imagenet classification,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1026–1034
2015
-
[62]
Improving the Capacity of Very Deep Networks with Maxout Units,
O. K. Oyedotun, A. E. R. Shabayek, D. Aouada, and B. Ottersten, “Improving the Capacity of Very Deep Networks with Maxout Units,” in ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2018
2018
-
[63]
FractalNet: Ultra-Deep Neural Networks without Residuals,
G. Larsson, M. Maire, and G. Shakhnarovich, “FractalNet: Ultra-Deep Neural Networks without Residuals,” in International Conference on Learning Representations (ICLR), 2017
2017
-
[64]
Shake-Shake regularization of 3-branch residual networks,
X. Gastaldi, “Shake-Shake regularization of 3-branch residual networks,” in International Conference on Learning Representations (ICLR) Work- shop, 2017
2017
-
[65]
Wide Residual Networks,
S. Zagoruyko and N. Komodakis, “Wide Residual Networks,” in British Machine Vision Conference (BMVC), 2016
2016
-
[66]
Designing Neural Network Architectures Using Reinforcement Learning,
B. Baker, O. Gupta, N. Naik, and R. Raskar, “Designing Neural Network Architectures Using Reinforcement Learning,” in Proceedings of the 5th International Conference on Learning Representations (ICLR) , 2017, pp. 1–18
2017
-
[67]
NSGA-NET: A Multi-Objective Genetic Algorithm for Neu- ral Architecture Search,
Z. Lu et al., “NSGA-NET: A Multi-Objective Genetic Algorithm for Neu- ral Architecture Search,” in The Genetic and Evolutionary Computation Conference (GECCO), 2019
2019
-
[68]
Neural Architecture Optimization,
R. Luo, F. Tian, T. Qin, E. Chen, and T.-Y . Liu, “Neural Architecture Optimization,” pp. 7816–7827, 2018. VOLUME 9, 2016 15 Muravev et al.: Neural Architecture Search by Estimation of Network Structure Distributions
2018
-
[69]
Hyperband: A Novel Bandit-Based Approach to Hyperparameter Opti- mization,
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Opti- mization,” Journal of Machine Learning Research , vol. 18, no. 185, pp. 1–52, 2018
2018
-
[70]
Cyclic Differentiable Architecture Search,
H. Yu and H. Peng, “Cyclic Differentiable Architecture Search,” arXiv preprint arXiv:2006.10724, 6 2020. ANTON MURAVEVreceived his B.Sc. and M.Sc. degrees in computer science from the Tomsk Poly- technic University, Tomsk, Russia, followed by the M.Sc. degree in information te...
2006 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.