Pith. sign in

REVIEW 4 major objections 6 minor 70 references

Neural Architecture Search by Estimation of Network Structure Distributions

T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Irregular neural network architectures can be found with a probability matrix alone.

desk verdict Novel representation and honest limitations, but the EDA update's value over random sampling is unshown; worth refereeing, not citing yet. read the letter →

arxiv 1908.06886 v3 pith:3YXVXCG3 submitted 2019-08-19 cs.NE cs.AIcs.LG

classification cs.NEcs.AIcs.LG
keywords neuralarchitecturesearchestimationofdistributionalgorithmsprobabilisticnetworkrepresentationconvolutionalnetworksCIFAR-100USPSwithoutcells
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces ASED, a neural architecture search method that represents a whole feedforward convolutional network as a matrix of layer-type probabilities instead of as a stack of repeated cells or blocks. Each row of the matrix is an independent categorical distribution over ten operations, and the search iterates by sampling networks, briefly training them, re-estimating the matrix from the best candidates, and appending new rows to grow deeper. The paper claims that this representation can reach irregular architectures that cell-based spaces cannot express, and that the discovered models are competitive in accuracy and computational cost. On USPS, the found architectures roughly match a heavily modified ResNet with far fewer parameters. On CIFAR-100, the best ASED model reaches 77.3% accuracy with 16.9 million parameters after 20 GPU-days of search.

What carries the argument

The prototype matrix $P$ is the load-bearing object: a discrete probability distribution over layer types for every position in a growing feedforward network. The update rule (Eq. 2) re-estimates $P$ from the empirical layer choices of the best $K_s$ sampled networks, which is the same marginal-re-estimation step used by univariate estimation of distribution algorithms. Around this core, the paper adds three mechanisms: probability capping (clamping row entries to $[p_{\min}, p_{\max}]$), prototype inversion (replacing high probabilities with low ones when the mean row $L^2$-norm crosses a threshold, to escape premature convergence), and fixed shortcut patterns—residual or semi-dense—that are applied to every sampled network so deeper candidates train more reliably.

What would settle it

Take a sample of architectures from a late-stage prototype, train each for the 20-epoch brief regime and again for the 200-epoch full regime, and compute the rank correlation between the two accuracy orderings; if the correlation is near zero, the search's selection signal is dominated by training-speed artifacts rather than final model quality.

Watch

Extended reading notes

Core claim

The paper's central claim is that a single probability matrix—a prototype—can stand in for an entire population of network architectures and can be optimized to produce useful, non-regular CNNs. Under the assumption that each layer type is chosen independently, the prototype $P$ has one row per current layer and one column per operation in the layer library; sampling a network from $P$ gives a concrete feedforward architecture. After $K$ candidates are briefly trained and ranked, the top $K_s$ are used to set $P_{ij}=\frac{1}{|K_s|}\sum_{k=1}^{K_s} x^k_{ij}$, the empirical frequency of operation $j$ at layer $i$ among the selected models, and this is followed by appending newly initialized rows. The authors show that the resulting search discovers architectures without repeating operation sequences, such as a CIFAR-100 net dominated by large 7x7 and dilated 5x5 convolutions in later layers, and they report that these architectures are competitive with existing methods even though the search uses no weight sharing and only 20-epoch candidate training.

Load-bearing premise

The search assumes that the ranking of candidates after only 20 training epochs is accurate enough to guide re-estimation, and the paper concedes in Section IV-C that this condition does not strictly hold, with the baseline variant scoring best under brief training while ultimately performing worst.

Editorial extensions

If this is right

  • The independence assumption plus a simple sampling loop is enough to find competitive CNNs, so the search space itself—not gradient-based architecture optimization or reinforcement learning—can carry much of the work.
  • Irregular architectures that mix large kernels, dilated convolutions, and pooling at varying depths are reachable and can be competitive, so cell-based spaces may be leaving useful designs undiscovered.
  • Because candidate networks are trained independently, the search parallelizes almost linearly with the number of GPUs.
  • The algorithm's low-complexity bias means it naturally favors simpler, faster-to-train models, which is useful when a compact deployment target is more important than peak accuracy.
  • Probability capping can stall the search, while inversion and shortcuts improve final accuracy, showing that convergence-control choices materially change the outcome.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The brief-training rank distortion the authors document suggests a promising extension: replacing the fixed 20-epoch evaluation with adaptive budgets, such as spending more epochs on promising candidates, could improve final architectures without scaling cost linearly.
  • Prototype inversion behaves like a tabu-style diversification mechanism; one could test whether inverting only the most certain rows (partial inversion) versus all rows is better on deeper searches, especially when combined with shortcut patterns.
  • The same prototype representation could be applied to other structure-selection tasks, such as choosing operations in recurrent cells or transformer layers, where the independence assumption is even more approximate but may still provide a useful prior.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper proposes ASED, an architecture search method that represents a CNN by a matrix of independent per-layer probabilities over ten layer types, updates the matrix with a UMDA-style re-estimation from the best Ks of K briefly trained candidates, and progressively adds layers. It reports experiments on USPS and CIFAR-100, comparing variants with probability capping, prototype inversion, and residual/semi-dense shortcut patterns, and it claims to discover irregular architectures competitive in accuracy and compute with cell-based NAS methods. The paper also includes an initialization baseline and a random-uniform architecture baseline.

Significance. If the empirical claims hold, ASED makes a useful contribution by demonstrating that a simple probability-matrix representation can search a non-cell, continuously growing architecture space, with interpretable state and easy parallelization. The authors are honest about limitations and provide source code. The reported best architectures exhibit genuinely non-repeating layer patterns that cannot be represented by identical cells, which is valuable. However, the strength of the causal claim that the EDA update drives search improvement is currently not established, and several comparisons lack the statistical and procedural controls needed to justify the word 'competitive.'

major comments (4)
  1. [Section IV-C, Fig. 3, Table 2] The paper explicitly states in Section IV-C that the assumption that brief-training rankings match full-training rankings 'does not strictly hold,' and the baseline variant has the highest brief-training validation accuracy but the weakest final accuracy. Since Eq. (2) re-estimates the prototype from the top-Ks briefly trained candidates, the central mechanism—that the prototype is tuned toward high-performance models—is not supported unless the selection signal is validated. I request a quantitative analysis of the rank correlation between 20-epoch and 200-epoch validation performance for a sample of architectures, or, alternatively, a reframing of the contribution as a search heuristic whose improvements come from the inversion/shortcut perturbations rather than from the EDA update.
  2. [Table 2] Every ASED variant is reported as a single best architecture from one stochastic search, with no repeated runs, confidence intervals, or significance tests. At 256 channels the ASED baseline (0.7483) is essentially matched by a single sample of 1000 networks from a uniform 16-layer prototype (0.7499), so the result does not demonstrate that iterative prototype updates improve over random sampling. The full-inversion variant's 0.7729 is promising, but without multiple seeds it is not possible to rule out seed luck. Please report repeated independent searches (mean and standard deviation, or at least best-of-k with error bars), and compare against random sampling under a matched computational budget.
  3. [Table 3] The comparison to published NAS methods in Table 3 is not made under a unified training setup; the text acknowledges that the results were obtained under non-matching environments and that some numbers are borrowed from other papers (e.g., PNAS/ENAS/DARTS/NAONet). Because final accuracy is highly sensitive to training schedule, regularization, and preprocessing, the claim that ASED is 'competitive both in accuracy and computational cost' is not established by this table. Please evaluate the final ASED architecture under the same training protocol as at least one strong competitor, or use a standardized benchmark such as NAS-Bench, or restrict the competitive claim to the internal baselines and non-cell-based methods.
  4. [Algorithm 1 and Section IV-A] Algorithm 1 returns the prototype P and says the final architecture is the one with highest probability, but the experiments report the 'best discovered architectures' and the 'best-performing network structures from each algorithm variant.' These are different selection rules: choosing the highest-probability architecture is a model-based output, while choosing the best validation-scored sampled network is best-of-search and inflates performance via selection bias. The paper must specify which rule was used and, if best-of-search, analyze the validation-gap/selection bias or use a hold-out selection procedure.
minor comments (6)
  1. [Equation (4)] The normalization formula is incomplete as written; define all symbols and specify how the row-wise scaling is applied when multiple entries are capped.
  2. [Section IV-A] 'PReLu' should be 'PReLU'; also 'We choose the follow' is a typo for 'We choose to follow.'
  3. [Figure 3] The caption references red dotted and green dashed lines but the text should state which colors correspond to max and median, and the y-axis label is missing units (accuracy).
  4. [Table 3] The row 'ASED (best)' should identify the variant (full inversion), the channel count (256), and the source of the 20-GPU-day search-cost estimate.
  5. [Section II] The related-work discussion would benefit from a direct comparison of ASED to PARSEC in the experiments or at least a clear statement of why PARSEC was not included in Table 3.
  6. [Section III-D] The shortcut patterns are fixed rules rather than learned structures; this should be stated more prominently in the abstract or contributions so that readers do not over-interpret 'non-linear connectivity' as optimized connectivity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: ASED is an empirical EDA-style search with external baselines; the admitted brief-training ranking distortion is a validity threat, not a circular derivation.

full rationale

This paper contains no mathematical derivation whose conclusion equals its premises; it is an empirical EDA-style search procedure. The prototype update in Eq. (2) is exactly the intended UMDA-style selection step, and the final architecture is read off the converged prototype; nothing is 'predicted' from a fitted quantity in the circular sense. The paper provides external baselines (best architecture from the initialization sample and best architecture from a uniformly random 16-layer prototype in Table 2), so the central comparison is falsifiable rather than constructed. The admitted failure of the brief-training ranking assumption in Section IV-C ('this condition does not strictly hold... validation results can be confusing for the search') is a threat to the validity of the selection signal and to the strength of the central claim, but it is an empirical assumption about training fidelity, not a definitional or self-citational circle. Self-citations (e.g., [12], [22], [45]-[47]) appear only as background and are not load-bearing for the ASED claim. No uniqueness theorem, ansatz smuggled through citation, or renamed known result was found. Although the random-uniform baseline nearly matches the baseline ASED variant at 256 channels, that is a weakness of the mechanism's demonstrated effect, not circularity.

Assumptions & free parameters 5 free parameters · 3 assumptions · 0 invented entities

The central claim relies on a small set of algorithm hyperparameters (listed above) and three domain assumptions: independence of layer choices, a fixed layer library with bound hyperparameters, and the validity of brief-training rankings. The brief-training ranking assumption is the most fragile, as the paper itself shows it is not strictly met. No new physical or mathematical entities are introduced.

free parameters (5)
  • Ninit = 5 (default), 2 (modified)
    Starting prototype depth, chosen by hand. A small value risks premature convergence, a large value makes search harder (Section IV-A).
  • K and Ks = K=1000, Ks=100 (default); K=100, Ks=10 (modified)
    Number of sampled candidates and selected top candidates per iteration; empirically set.
  • pmax = 0.9
    Upper probability cap for probability capping; pmin is computed from pmax using Eq. 3.
  • L2-norm inversion threshold = 0.65
    Threshold on the mean row L2-norm to trigger prototype inversion; set to the middle of the possible interval.
  • Shortcut parameter D = 2 or 3
    Controls residual and semi-dense shortcut patterns; D=1 has limited impact and D>3 creates too few shortcuts.
assumptions (3)
  • domain assumption Layer type choices are independent across positions (Section III-A).
    The paper states this is 'unlikely to hold in practice' but is used for tractability, with inter-layer interactions implicitly handled during search.
  • domain assumption Only layer types are optimized; filter sizes, strides, and channel counts are fixed (Section III-A).
    The search space is limited to discrete layer type choices, with all other hyperparameters bound to the layer type or set to a constant (32 channels).
  • domain assumption Brief training (20 epochs) produces rankings that correlate with final performance (Section IV-A, IV-C).
    The algorithm's prototype updates depend on these rankings, but the authors demonstrate the assumption is partially violated, causing search distortion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Neural Architecture Search by Estimation of Network Structure Distributions." pith.science (2026). https://pith.science/paper/3YXVXCG3

@misc{pith2026190806886,
  author       = {Pith},
  title        = {Pith review of: Neural Architecture Search by Estimation of Network Structure Distributions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3YXVXCG3}},
  note         = {Machine review of arXiv:1908.06886}
}
read the original abstract

The influence of deep learning is continuously expanding across different domains, and its new applications are ubiquitous. The question of neural network design thus increases in importance, as traditional empirical approaches are reaching their limits. Manual design of network architectures from scratch relies heavily on trial and error, while using existing pretrained models can introduce redundancies or vulnerabilities. Automated neural architecture design is able to overcome these problems, but the most successful algorithms operate on significantly constrained design spaces, assuming the target network to consist of identical repeating blocks. While such approach allows for faster search, it does so at the cost of expressivity. We instead propose an alternative probabilistic representation of a whole neural network structure under the assumption of independence between layer types. Our matrix of probabilities is equivalent to the population of models, but allows for discovery of structural irregularities, while being simple to interpret and analyze. We construct an architecture search algorithm, inspired by the estimation of distribution algorithms, to take advantage of this representation. The probability matrix is tuned towards generating high-performance models by repeatedly sampling the architectures and evaluating the corresponding networks, while gradually increasing the model depth. Our algorithm is shown to discover non-regular models which cannot be expressed via blocks, but are competitive both in accuracy and computational cost, while not utilizing complex dataflows or advanced training techniques, as well as remaining conceptually simple and highly extensible.

Figures

Figures reproduced from arXiv: 1908.06886 by the authors.

Figure 1
Figure 1. FIGURE 1 [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. FIGURE 2 [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. FIGURE 3 [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: FIGURE 4 [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 67 canonical work pages

  1. [1]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 5 2015

  2. [2]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778

  3. [3]

    Densely connected convolutional networks,

    G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2261–2269

  4. [4]

    DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,

    L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 4, pp. 834–848, 2018

  5. [5]

    Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems 28, pp. 91–99, 2015

  6. [6]

    The Mythos of Model Interpretability,

    Z. C. Lipton, “The Mythos of Model Interpretability,” in ICML 2016 Workshop on Human Interpretability in Machine Learning (WHI 2016) , 2016

  7. [7]

    Interpreting Deep Learning: The Machine Learning Rorschach Test?

    A. S. Charles, “Interpreting Deep Learning: The Machine Learning Rorschach Test?”arXiv preprint arXiv:1806.00148, 2018

  8. [8]

    Methods for interpreting and understanding deep neural networks,

    G. Montavon, W. Samek, and K.-R. Müller, “Methods for interpreting and understanding deep neural networks,” Digital Signal Processing, vol. 73, pp. 1–15, 2018

Show all 70 references
  1. [9]

    Deep Neural Networks Are Easily Fooled: High Confidence Predictions for Unrecognizable Images,

    A. Nguyen, J. Yosinski, and J. Clune, “Deep Neural Networks Are Easily Fooled: High Confidence Predictions for Unrecognizable Images,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 427–436

  2. [10]

    Identity mappings in deep residual networks,

    K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision (ECCV) , 2016, pp. 630–645

  3. [11]

    A survey of transfer learning,

    K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,”Journal of Big Data, vol. 3, no. 1, p. 9, 2016

  4. [12]

    On the Layer Selection in Small-Scale Deep Networks,

    A. Muravev, J. Raitoharju, and M. Gabbouj, “On the Layer Selection in Small-Scale Deep Networks,” in 7th European Workshop on Visual Information Processing (EUVIP), 2018

  5. [13]

    A Deeper Look at Dataset Bias,

    T. Tommasi, N. Patricia, B. Caputo, and T. Tuytelaars, “A Deeper Look at Dataset Bias,” in Domain Adaptation in Computer Vision Applications, G. Csurka, Ed. Cham: Springer International Publishing, 2017, pp. 37– 55

  6. [14]

    Genetic algorithms and neural networks: optimizing connections and connectivity,

    D. Whitley, T. Starkweather, and C. Bogart, “Genetic algorithms and neural networks: optimizing connections and connectivity,”Parallel Com- puting, vol. 14, no. 3, pp. 347–361, 1990

  7. [15]

    Evolving artificial neural networks,

    Xin Yao, “Evolving artificial neural networks,” Proceedings of the IEEE, vol. 87, no. 9, pp. 1423–1447, 1999

  8. [16]

    Evolving Neural Network through Augmenting Topologies,

    K. O. Stanley and R. Miikkulainen, “Evolving Neural Network through Augmenting Topologies,” Evolutionary Computation, vol. 10, no. 2, pp. 99–127, 2002

  9. [17]

    A Hypercube-Based Encoding for Evolving Large-Scale Neural Networks,

    K. O. Stanley, D. B. D’Ambrosio, and J. Gauci, “A Hypercube-Based Encoding for Evolving Large-Scale Neural Networks,” Artificial Life , vol. 15, no. 2, pp. 185–212, 2009. 14 VOLUME 9, 2016 Muravev et al.: Neural Architecture Search by Estimation of Network Structure Distributions

  10. [18]

    Evolutionary artifi- cial neural networks by multi-dimensional particle swarm optimization,

    S. Kiranyaz, T. Ince, A. Yildirim, and M. Gabbouj, “Evolutionary artifi- cial neural networks by multi-dimensional particle swarm optimization,” Neural Networks, vol. 22, no. 10, pp. 1448–1462, 2009

  11. [19]

    Neuroevolution: from architec- tures to learning,

    D. Floreano, P. Dürr, and C. Mattiussi, “Neuroevolution: from architec- tures to learning,”Evolutionary Intelligence, vol. 1, no. 1, pp. 47–62, 2008

  12. [20]

    Genetic CNN,

    L. Xie and A. Yuille, “Genetic CNN,” in IEEE International Conference on Computer Vision (ICCV), 2017, pp. 1388–1397

  13. [21]

    Large-Scale Evolution of Image Classifiers,

    E. Real et al., “Large-Scale Evolution of Image Classifiers,” in Interna- tional Conference on Machine Learning (ICML), 2017, pp. 2902–2911

  14. [22]

    Finding Better Topologies for Deep Convolutional Neural Networks by Evolution,

    H. Zhang, S. Kiranyaz, and M. Gabbouj, “Finding Better Topologies for Deep Convolutional Neural Networks by Evolution,” arXiv preprint arXiv:1809.03242, 2018

  15. [23]

    Regularized Evolution for Image Classifier Architecture Search,

    E. Real, A. Aggarwal, Y . Huang, and Q. V . Le, “Regularized Evolution for Image Classifier Architecture Search,” Proceedings of the AAAI Confer- ence on Artificial Intelligence, vol. 33, pp. 4780–4789, 2019

  16. [24]

    Evolving Deep Convolutional Neural Networks for Image Classification,

    Y . Sun, B. Xue, M. Zhang, and G. G. Yen, “Evolving Deep Convolutional Neural Networks for Image Classification,” IEEE Transactions on Evolu- tionary Computation, vol. 24, pp. 394–407, 2019

  17. [25]

    Neural Architecture Search with Reinforce- ment Learning,

    B. Zoph and Q. V . Le, “Neural Architecture Search with Reinforce- ment Learning,” inInternational Conference on Learning Representations (ICLR), 2017

  18. [26]

    Learning Transferable Architectures for Scalable Image Recognition,

    B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning Transferable Architectures for Scalable Image Recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  19. [27]

    Efficient Neural Architecture Search via Parameter Sharing,

    H. Pham, M. Y . Guan, B. Zoph, Q. V . Le, and J. Dean, “Efficient Neural Architecture Search via Parameter Sharing,” in Proceedings of the 35th International Conference on Machine Learning (PMLR), 2018, pp. 4095– 4104

  20. [28]

    Progressive Neural Architecture Search,

    C. Liu et al. , “Progressive Neural Architecture Search,” in European Conference on Computer Vision (ECCV). Cham: Springer International Publishing, 2018, pp. 19–35

  21. [29]

    Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution,

    T. Elsken, J. H. Metzen, and F. Hutter, “Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution,” in International Confer- ence on Learning Representations (ICLR), 2019

  22. [30]

    DARTS: Differentiable Architecture Search,

    H. Liu, K. Simonyan, and Y . Yang, “DARTS: Differentiable Architecture Search,” inInternational Conference on Learning Representations (ICLR), 2019

  23. [31]

    Evaluat- ing the Search Phase of Neural Architecture Search,

    C. Sciuto, K. Yu, M. Jaggi, C. Musat, and M. Salzmann, “Evaluat- ing the Search Phase of Neural Architecture Search,” arXiv preprint arXiv:1902.08142, 2019

  24. [32]

    Population-Based Incremental Learning: A Method for In- tegrating Genetic Search Based Function Optimization and Competitive Learning,

    S. Baluja, “Population-Based Incremental Learning: A Method for In- tegrating Genetic Search Based Function Optimization and Competitive Learning,” Pittsburgh, PA, USA, 1994

  25. [33]

    The Equation for Response to Selection and Its Use for Prediction,

    H. Mühlenbein, “The Equation for Response to Selection and Its Use for Prediction,”Evolutionary Computation, vol. 5, no. 3, pp. 303–346, 1997

  26. [34]

    Learning Represen- tations by Back-propagating Errors,

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning Represen- tations by Back-propagating Errors,” in Neurocomputing: Foundations of Research, J. A. Anderson and E. Rosenfeld, Eds. Cambridge, MA, USA: MIT Press, 1988, pp. 696–699

  27. [35]

    Evolutionary computation: comments on the history and current state,

    T. Back, U. Hammel, and H.-P. Schwefel, “Evolutionary computation: comments on the history and current state,” IEEE Transactions on Evo- lutionary Computation, vol. 1, no. 1, pp. 3–17, 1997

  28. [36]

    Luke, Essentials of Metaheuristics, 2nd ed

    S. Luke, Essentials of Metaheuristics, 2nd ed. Lulu, 2013

  29. [37]

    Incremental Evolution of Complex General Behavior,

    F. Gomez and R. Miikkulainen, “Incremental Evolution of Complex General Behavior,”Adaptive Behavior, vol. 5, no. 3-4, pp. 317–342, 1997

  30. [38]

    A new evolutionary system for evolving artificial neural networks,

    X. Yao and Y . Liu, “A new evolutionary system for evolving artificial neural networks,” IEEE Transactions on Neural Networks , vol. 8, no. 3, pp. 694–713, 5 1997

  31. [39]

    Solving non-Markovian Control Tasks with Neuroevolution,

    F. J. Gomez and R. Miikkulainen, “Solving non-Markovian Control Tasks with Neuroevolution,” in Proceedings of the 16th International Joint Conference on Artificial Intelligence - Volume 2, 1999, pp. 1356–1361

  32. [40]

    Compositional pattern producing networks: A novel abstraction of development,

    K. O. Stanley, “Compositional pattern producing networks: A novel abstraction of development,” Genetic Programming and Evolvable Ma- chines, vol. 8, no. 2, pp. 131–162, 2007

  33. [41]

    An Enhanced Hypercube-Based Encoding for Evolving the Placement, Density, and Connectivity of Neurons,

    S. Risi and K. O. Stanley, “An Enhanced Hypercube-Based Encoding for Evolving the Placement, Density, and Connectivity of Neurons,”Artificial Life, vol. 18, no. 4, pp. 331–363, 2012

  34. [42]

    HyperNEAT: The First Five Years,

    D. B. D’Ambrosio, J. Gauci, and K. O. Stanley, “HyperNEAT: The First Five Years,” in Growing Adaptive Machines . Springer, Berlin, Heidelberg, 2014, pp. 159–185

  35. [43]

    Evolving Deep Neural Networks,

    R. Miikkulainen et al., “Evolving Deep Neural Networks,” inArtificial In- telligence in the Age of Neural Networks and Brain Computing. Academic Press, 2019, pp. 293 – 312

  36. [44]

    Completely Automated CNN Architecture Design Based on Blocks,

    Y . Sun, B. Xue, M. Zhang, and G. G. Yen, “Completely Automated CNN Architecture Design Based on Blocks,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–13, 2019

  37. [45]

    Progressive Opera- tional Perceptrons,

    S. Kiranyaz, T. Ince, A. Iosifidis, and M. Gabbouj, “Progressive Opera- tional Perceptrons,”Neurocomputing, vol. 224, pp. 142–154, 2017

  38. [46]

    Operational Neural Networks,

    ——, “Operational Neural Networks,” Neural Computing and Applica- tions (in print), 2020

  39. [47]

    Heterogeneous Multilayer Generalized Operational Perceptron,

    D. T. Tran, S. Kiranyaz, M. Gabbouj, and A. Iosifidis, “Heterogeneous Multilayer Generalized Operational Perceptron,” IEEE Transactions on Neural Networks and Learning Systems, pp. 1–15, 2019

  40. [48]

    SMASH: One-Shot Model Architecture Search through HyperNetworks,

    A. Brock, T. Lim, J. M. Ritchie, and N. Weston, “SMASH: One-Shot Model Architecture Search through HyperNetworks,” in Workshop on Meta-Learning (MetaLearn 2017) at NIPS, 2017

  41. [49]

    Simple and Efficient Architecture Search for Convolutional Neural Networks,

    T. Elsken, J.-H. Metzen, and F. Hutter, “Simple and Efficient Architecture Search for Convolutional Neural Networks,” in 6th International Confer- ence on Learning Representations (ICLR), 2018

  42. [50]

    In- staNAS: Instance-aware Neural Architecture Search,

    A.-C. Cheng, C. H. Lin, D.-C. Juan, W. Wei, and M. Sun, “In- staNAS: Instance-aware Neural Architecture Search,” arXiv preprint arXiv:1811.10201, 2018

  43. [51]

    Neural Architecture Search with Bayesian Optimisation and Optimal Transport,

    K. Kandasamy, W. Neiswanger, J. Schneider, B. Poczos, and E. P. Xing, “Neural Architecture Search with Bayesian Optimisation and Optimal Transport,” in Advances in Neural Information Processing Systems 31 , 2018, pp. 2016–2025

  44. [52]

    Probabilistic Neural Architecture Search,

    F. P. Casale, J. Gordon, and N. Fusi, “Probabilistic Neural Architecture Search,”arXiv preprint arXiv:1902.05116, 2019

  45. [53]

    A Survey of Optimization by Building and Using Probabilistic Models,

    M. Pelikan, D. E. Goldberg, and F. G. Lobo, “A Survey of Optimization by Building and Using Probabilistic Models,” Computational Optimization and Applications, vol. 21, no. 1, pp. 5–20, 2002

  46. [54]

    An introduction and survey of estimation of distribution algorithms,

    M. Hauschild and M. Pelikan, “An introduction and survey of estimation of distribution algorithms,”Swarm and Evolutionary Computation, vol. 1, no. 3, pp. 111–128, 2011

  47. [55]

    Comparing two K-category assignments by a K-category correlation coefficient,

    J. Gorodkin, “Comparing two K-category assignments by a K-category correlation coefficient,” Computational Biology and Chemistry , vol. 28, no. 5-6, pp. 367–374, 12 2004

  48. [56]

    On Stability of Fixed Points of Limit Models of Univariate Marginal Distribution Algorithm and Factorized Distribution Algorithm,

    Q. Zhang, “On Stability of Fixed Points of Limit Models of Univariate Marginal Distribution Algorithm and Factorized Distribution Algorithm,” IEEE Transactions on Evolutionary Computation, vol. 8, no. 1, pp. 80–93, 2004

  49. [57]

    EDAs Cannot Be Balanced and Stable,

    T. Friedrich, T. Kötzing, and M. S. Krejca, “EDAs Cannot Be Balanced and Stable,” inProceedings of the Genetic and Evolutionary Computation Conference, 2016, pp. 1139–1146

  50. [58]

    Tabu Search,

    F. Glover and M. Laguna, “Tabu Search,” in Handbook of Combinatorial Optimization, D.-Z. Du and P. M. Pardalos, Eds. Boston, MA: Springer US, 1998, pp. 2093–2229

  51. [59]

    A Database for Handwritten Text Recognition Research,

    J. J. Hull, “A Database for Handwritten Text Recognition Research,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 1994

  52. [60]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009

  53. [61]

    Delving deep into rectifiers: Surpass- ing human-level performance on imagenet classification,

    K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpass- ing human-level performance on imagenet classification,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1026–1034

  54. [62]

    Improving the Capacity of Very Deep Networks with Maxout Units,

    O. K. Oyedotun, A. E. R. Shabayek, D. Aouada, and B. Ottersten, “Improving the Capacity of Very Deep Networks with Maxout Units,” in ICASSP , IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2018

  55. [63]

    FractalNet: Ultra-Deep Neural Networks without Residuals,

    G. Larsson, M. Maire, and G. Shakhnarovich, “FractalNet: Ultra-Deep Neural Networks without Residuals,” in International Conference on Learning Representations (ICLR), 2017

  56. [64]

    Shake-Shake regularization of 3-branch residual networks,

    X. Gastaldi, “Shake-Shake regularization of 3-branch residual networks,” in International Conference on Learning Representations (ICLR) Work- shop, 2017

  57. [65]

    Wide Residual Networks,

    S. Zagoruyko and N. Komodakis, “Wide Residual Networks,” in British Machine Vision Conference (BMVC), 2016

  58. [66]

    Designing Neural Network Architectures Using Reinforcement Learning,

    B. Baker, O. Gupta, N. Naik, and R. Raskar, “Designing Neural Network Architectures Using Reinforcement Learning,” in Proceedings of the 5th International Conference on Learning Representations (ICLR) , 2017, pp. 1–18

  59. [67]

    NSGA-NET: A Multi-Objective Genetic Algorithm for Neu- ral Architecture Search,

    Z. Lu et al., “NSGA-NET: A Multi-Objective Genetic Algorithm for Neu- ral Architecture Search,” in The Genetic and Evolutionary Computation Conference (GECCO), 2019

  60. [68]

    Neural Architecture Optimization,

    R. Luo, F. Tian, T. Qin, E. Chen, and T.-Y . Liu, “Neural Architecture Optimization,” pp. 7816–7827, 2018. VOLUME 9, 2016 15 Muravev et al.: Neural Architecture Search by Estimation of Network Structure Distributions

  61. [69]

    Hyperband: A Novel Bandit-Based Approach to Hyperparameter Opti- mization,

    L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Opti- mization,” Journal of Machine Learning Research , vol. 18, no. 185, pp. 1–52, 2018

  62. [70]

    Cyclic Differentiable Architecture Search,

    H. Yu and H. Peng, “Cyclic Differentiable Architecture Search,” arXiv preprint arXiv:2006.10724, 6 2020. ANTON MURAVEVreceived his B.Sc. and M.Sc. degrees in computer science from the Tomsk Poly- technic University, Tomsk, Russia, followed by the M.Sc. degree in information te...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.