REVIEW 3 major objections 6 minor 31 references
Efficient and Diverse Generative Robot Designs using Evolution and Intrinsic Motivation
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Homeokinesis makes evolved robot designs outperform fixed-controller baselines on every tested task.
desk verdict Homeokinesis as a cheap evaluation proxy in morpho-evolution is a real idea, but the downstream superiority claim is confounded by asymmetric selection thresholds and a missing MEL baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the homeokinetic controller, a pseudo-linear controller updated by stochastic gradient descent on the time-loop error $\mathrm{TLE} = \| L^{-1} E \|^2$, where $E$ is the forward-model prediction error and $L$ the Jacobian of the sensorimotor loop. Minimising the TLE balances predictable behaviour against sensitive behaviour, producing exploratory motion in a few seconds with no task reward. In MEHK this controller replaces the learning phase: each design's fitness is the fraction of 64 arena cells visited during 20 minutes of homeokinetic exploration, and evolution runs through the asynchronous AME algorithm on CPPN-encoded bodies.
What would settle it
Evolve designs with MEHK but replace the homeokinetic score with a random score that matches the same distribution of exploration values, keeping the Pareto selection and NCMA-ES training identical; if downstream performance does not drop, the homeokinetic signal itself is not carrying the result.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that homeokinesis, a parameter-free intrinsic-motivation controller that maximises predictability and sensitivity via the time-loop error, is an effective and cheap evaluation policy for morpho-evolution. By scoring each robot body on how much of an 8-by-8 arena it covers in 20 minutes of homeokinetic exploration, the evolutionary loop can filter for functionally viable bodies—ones that move and interact with the environment—without task-specific training. From 10,000 generated designs, the authors select three per replicate from a Pareto front of exploration fitness and morphological sparsity, train Elman-network controllers with NCMA-ES, and find that the MEHK bodies outperform MEFC bodies on all four downstream tasks, with statistical significance at p<$10^{-4}$. The method generates 10,000 robots in about 40 hours on 64 CPUs, which the authors contrast with the 1,152 CPUs and 4,000 designs reported for a deep-RL morpho-evolution baseline.
Load-bearing premise
The 20-minute exploration score under homeokinesis is a valid proxy for whether a robot body will train well on unrelated later tasks; the authors themselves note the score has minimal impact on downstream performance.
Editorial extensions
If this is right
- MEHK designs outperform MEFC designs in all four downstream tasks, so the homeokinetic exploration score is a viable cheap proxy for body viability in morpho-evolution.
- The method generates and evaluates 10,000 robot designs in about 40 hours on 64 CPUs, substantially lowering the compute barrier compared to deep-RL morpho-evolution approaches.
- MEHK produces a more diverse mixture of component types than MEFC, countering the premature convergence that learning-based evaluation tends to induce.
- The Pareto selection between exploration fitness and morphological sparsity yields both generalist robots and specialists that excel at single tasks.
Reading between the lines
- The paper's own observation that the exploration score has minimal impact on downstream performance suggests the advantage of MEHK may come primarily from the diversity of body morphologies it produces rather than from homeokinesis selecting intrinsically 'good' bodies; a selection mechanism preserving the same diversity might perform equally well.
- The same morpho-evolution framework could serve as a testbed for comparing other intrinsic-motivation signals, such as empowerment or predictive information, as cheap viability filters for robot bodies.
- The fixed-controller baseline's random weights bias it toward jointed bodies, so part of MEFC's weakness may stem from a controller-design mismatch rather than the absence of intrinsic motivation; a stronger baseline would use a controller tuned per morphology without full task-specific learning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes MEHK, a morpho-evolution algorithm that evaluates robot designs using a homeokinetic controller to generate exploratory behavior instead of training a task-specific controller during evolution. The authors compare MEHK to MEFC, a morpho-evolution baseline with a fixed feed-forward controller, on four downstream tasks after selecting three designs per replicate via a fitness threshold and a Pareto-based sparsity selection. They report that MEHK-generated designs achieve higher task scores on all four tasks, similar wall-clock time (about 40 hours for 10,000 designs on 64 CPUs), and different morphological diversity compared to MEFC. The manuscript includes 30 replicates per condition and reports Brunner-Munzel tests with p-values below 1e-4 on most tasks.
Significance. If the result holds, MEHK offers a computationally cheap method for generating robot morphologies without per-design learning, potentially scaling to larger design spaces. The paper is, to my knowledge, the first to combine homeokinesis with CPPN-based morpho-evolution and to include exteroceptive sensors in this framework. Strengths include the clean 30-replicate protocol, the use of statistical testing, and the analysis of morphological distributions. However, the central comparison is burdened by asymmetric selection thresholds between MEHK and MEFC, and the claimed comparison with 'current MEL methods' is not actually performed; the only baseline is a fixed-controller evolution. These issues require careful re-analysis and re-framing before the headline claims can be accepted.
major comments (3)
- [Section III-B (Robotic design selection)] The selection step uses asymmetric thresholds: '0.4 for MEHK and 0.2 for MEFC.' Since Fig. 5 shows MEHK exploration scores are systematically higher (median 0.8 vs 0.45), these thresholds select different percentiles of the two distributions. The paper's own Section IV-B states that exploration score has 'minimal impact' on downstream performance, which means the higher threshold does not have an obvious mechanism for selecting more trainable designs. The observed downstream advantage could therefore be caused by the threshold asymmetry or by morphological correlates (e.g., more wheels and fewer limbs in MEHK, per Fig. 6) rather than by homeokinesis acting as a viability filter. The authors should provide a threshold-sensitivity analysis with matched thresholds, or compare downstream performance of all generated designs without threshold filtering.
- [Abstract and Section I (Introduction)] The abstract claims comparison with 'current MEL methods,' but the only in-paper baseline is MEFC, a morpho-evolution with a fixed controller and no learning phase (Section II-A-c). This overstates the evidence: the efficiency argument relies on a comparison with prior work (Gupta et al. [6]) rather than an in-paper MEL baseline. Moreover, Fig. 5 shows MEHK and MEFC use the same wall-clock time, so the abstract's 'quickly generated compared to morpho-evolution with static parameters' is not supported by the reported experiment. Please either add a genuine MEL baseline (e.g., AME with CMA-ES evaluation) or carefully revise the claims and abstract to state that the comparison is against fixed-controller morpho-evolution.
- [Section I and Section IV-B] The paper motivates homeokinesis as a proxy for design viability: 'homeokinesis is a proxy towards optimal designs, filtering designs unsuitable for movement' (Section I) and 'The exploration task biases the designs for functional combinations...' (Section III-A). Yet Section IV-B states that 'the exploration score... has minimal impact on the performance of the downstream tasks.' The mechanism by which MEHK improves downstream performance is therefore not established, and the downstream advantage could be a side-effect of selection on morphological traits. I recommend reporting the correlation between exploration fitness and downstream task scores within each condition, and comparing random subsamples with matched thresholds, to directly test the proxy-validity claim.
minor comments (6)
- [Figure 2 caption] The caption contains a typo: 'homoekinesis' should be 'homeokinesis.'
- [Acknowledgements] The word 'Unver' should likely be 'University' in 'Edinburgh Napier Unver.'
- [Section III-B] The sparsity score uses 15 nearest neighbours; no justification or sensitivity analysis is provided for this choice.
- [Section IV-B] The hill-climbing advantage is modest (medians around 0.4 vs 0.5) and only a tail of MEHK robots performs substantially better; reporting effect sizes or confidence intervals alongside p-values would strengthen the conclusions.
- [Section IV-A] The statement that MEHK 'can generate a wider diversity of designs than MEFC when comparing components' needs reconciliation with Fig. 7, which shows MEFC has higher chassis diversity; the claims about diversity are therefore more nuanced than the text suggests.
- [Section III-C and reproducibility] The paper states 'The code will be available upon acceptance' but no code or detailed simulation parameters (physics engine, timestep, sensor noise model) are provided; for a computational paper, shipping code or a complete parameter listing would substantially improve reproducibility.
Circularity Check
No significant circularity: the downstream comparison is an independent empirical evaluation, not a reduction to inputs; the asymmetric selection threshold is a correctness confound, not a circular step.
full rationale
The derivation chain is self-contained and does not reduce to its own inputs. The paper evolves designs using a homeokinesis-driven exploration score, selects a small subset, and then independently trains those designs on four downstream tasks with NCMA-ES. The downstream task scores are never used during evolution, selection, or the homeokinetic fitness computation, so no fitted parameter is renamed as a prediction. The stated claim that 'the exploration score has minimal impact on the performance of the downstream tasks' (Section IV-B) further confirms that the downstream results are not a relabeling of the selection fitness. The self-citations to AME, NCMA-ES, ARE, and homeokinetic learning are prior building blocks with external published content; invoking them as components of the pipeline is not load-bearing circularity because the novelty and the reported comparison are empirical outcomes of the combined system. The asymmetric selection thresholds of 0.4 for MEHK and 0.2 for MEFC (Section III-B) are a legitimate methodological concern and could confound the comparison, but confounding is not circularity: the selected MEHK designs are still trained by a separate optimization process, and no downstream score is equal by construction to the exploration fitness. The paper is therefore best assessed as having no significant circularity, with the threshold asymmetry flagged as a correctness risk rather than as a circular step.
Assumptions & free parameters
free parameters (4)
- exploration fitness selection threshold =
0.4 for MEHK, 0.2 for MEFC
- novelty weight increment in NCMA-ES =
0.05
- number of nearest neighbours in sparsity score =
15
- hidden neurons in controller networks =
6
assumptions (4)
- domain assumption Homeokinesis as described in Der and Martius produces effective exploratory behavior in arbitrary robot morphologies within seconds.
- ad hoc to paper The coverage score over an 8 by 8 arena is a sufficient proxy for design viability on unseen downstream tasks.
- standard math The inverse of the Jacobian L exists for the sensorimotor loop, as required by the time-loop-error update in Eq. 1.
- ad hoc to paper The fixed-controller baseline MEFC is representative of 'morpho-evolution with static parameters' and is an adequate stand-in for the claimed comparison against MEL.
Cite this review
Pith. "Pith review of Efficient and Diverse Generative Robot Designs using Evolution and Intrinsic Motivation." pith.science (2026). https://pith.science/paper/IQZTR6AP
@misc{pith2026241118423,
author = {Pith},
title = {Pith review of: Efficient and Diverse Generative Robot Designs using Evolution and Intrinsic Motivation},
year = {2026},
howpublished = {\url{https://pith.science/paper/IQZTR6AP}},
note = {Machine review of arXiv:2411.18423}
}
read the original abstract
Methods for generative design of robot physical configurations can automatically find optimal and innovative solutions for challenging tasks in complex environments. The vast search-space includes the physical design-space and the controller parameter-space, making it a challenging problem in machine learning and optimisation in general. Evolutionary algorithms (EAs) have shown promising results in generating robot designs via gradient-free optimisation. Morpho-evolution with learning (MEL) uses EAs to concurrently generate robot designs and learn the optimal parameters of the controllers. Two main issues prevent MEL from scaling to higher complexity tasks: computational cost and premature convergence to sub-optimal designs. To address these issues, we propose combining morpho-evolution with intrinsic motivations. Intrinsically motivated behaviour arises from embodiment and simple learning rules without external guidance. We use a homeokinetic controller that generates exploratory behaviour in a few seconds with reduced knowledge of the robot's design. Homeokinesis replaces costly learning phases, reducing computational time and favouring diversity, preventing premature convergence. We compare our approach with current MEL methods in several downstream tasks. The generated designs score higher in all the tasks, are more diverse, and are quickly generated compared to morpho-evolution with static parameters.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[6]
Embodied intel- ligence via learning and evolution,
A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei, “Embodied intel- ligence via learning and evolution,” Nature communications, vol. 12, no. 1, p. 5721, 2021
2021
-
[1]
Art and the science of generative ai,
Z. Epstein, A. Hertzmann, I. of Human Creativity, M. Akten, H. Farid, J. Fjeld, M. R. Frank, M. Groh, L. Herman, N. Leach, et al., “Art and the science of generative ai,” Science, vol. 380, no. 6650, pp. 1110– 1111, 2023
work page 2023
-
[2]
Generative design: an explorative study,
F. Buonamici, M. Carfagni, R. Furferi, Y . V olpe, L. Governi, et al. , “Generative design: an explorative study,”Computer-Aided Design and Applications, vol. 18, no. 1, pp. 144–155, 2020
work page 2020
-
[3]
Data-efficient co-adaptation of morphology and behaviour with deep reinforcement learning,
K. S. Luck, H. B. Amor, and R. Calandra, “Data-efficient co-adaptation of morphology and behaviour with deep reinforcement learning,” in Conference on Robot Learning . PMLR, 2020, pp. 854–869
work page 2020
-
[4]
Reinforcement learning for freeform robot design,
M. Li, D. Matthews, and S. Kriegman, “Reinforcement learning for freeform robot design,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 8799–8806
work page 2024
-
[5]
Efficient automatic design of robots,
D. Matthews, A. Spielberg, D. Rus, S. Kriegman, and J. Bongard, “Efficient automatic design of robots,” Proceedings of the National Academy of Sciences , vol. 120, no. 41, p. e2305180120, 2023
work page 2023
-
[7]
Morpho evolution with learning using a controller archive as an inheritance mechanism,
L. K. Le Goff, E. Buchanan, E. Hart, A. E. Eiben, W. Li, M. De Carlo, A. F. Winfield, M. F. Hale, R. Woolley, M. Angus, et al. , “Morpho evolution with learning using a controller archive as an inheritance mechanism,” IEEE Transactions on Cognitive and Developmental Systems, vol. 15, no. 2, pp. 507–517, 2022
work page 2022
-
[8]
Lamarckian evolution of simulated modular robots,
M. Jelisavcic, K. Glette, E. Haasdijk, and A. Eiben, “Lamarckian evolution of simulated modular robots,” Frontiers in Robotics and AI , vol. 6, p. 9, 2019
work page 2019
Show all 31 references
-
[9]
Evaluation of frameworks that combine evolution and learning to design robots in complex morphological spaces,
W. Li, E. Buchanan, L. K. Le Goff, E. Hart, M. F. Hale, M. De Carlo, R. Wooley, A. F. Winfield, J. Timmis, and A. M. Tyrrell, “Evaluation of frameworks that combine evolution and learning to design robots in complex morphological spaces,” IEEE Transactions on Evolutionary Comp...
2023
-
[10]
Scalable co- optimization of morphology and control in embodied machines,
N. Cheney, J. Bongard, V . SunSpiral, and H. Lipson, “Scalable co- optimization of morphology and control in embodied machines,” Journal of The Royal Society Interface , vol. 15, no. 143, p. 20170937, 2018
2018
-
[11]
Investigating premature convergence in co- optimization of morphology and control in evolved virtual soft robots,
A. Mertan and N. Cheney, “Investigating premature convergence in co- optimization of morphology and control in evolved virtual soft robots,” in European Conference on Genetic Programming (Part of EvoStar) . Springer, 2024, pp. 38–55
2024
-
[12]
Real-world evolution adapts robot morphology and control to hard- ware limitations,
T. F. Nygaard, C. P. Martin, E. Samuelsen, J. Torresen, and K. Glette, “Real-world evolution adapts robot morphology and control to hard- ware limitations,” in Proceedings of the Genetic and Evolutionary Computation Conference, 2018, pp. 125–132
2018
-
[13]
Evolving 3d morphology and behavior by competition,
K. Sims, “Evolving 3d morphology and behavior by competition,” Artificial life, vol. 1, no. 4, pp. 353–372, 1994
1994
-
[14]
Automatic design and manufacture of robotic lifeforms,
H. Lipson and J. B. Pollack, “Automatic design and manufacture of robotic lifeforms,” Nature, vol. 406, no. 6799, pp. 974–978, 2000
2000
-
[15]
Unshackling evolution: evolving soft robots with multiple materials and a powerful generative encoding,
N. Cheney, R. MacCurdy, J. Clune, and H. Lipson, “Unshackling evolution: evolving soft robots with multiple materials and a powerful generative encoding,” ACM SIGEVOlution, vol. 7, no. 1, pp. 11–23, 2014
2014
-
[16]
If it evolves it needs to learn,
A. Eiben and E. Hart, “If it evolves it needs to learn,” in Proceed- ings of the 2020 Genetic and Evolutionary Computation Conference Companion, 2020, pp. 1383–1384
2020
-
[17]
The effects of learning in morphologically evolving robot systems,
J. Luo, A. C. Stuurman, J. M. Tomczak, J. Ellers, and A. E. Eiben, “The effects of learning in morphologically evolving robot systems,” Frontiers in Robotics and AI , vol. 9, p. 797393, 2022
2022
-
[18]
On the difficulty of co-optimizing morphology and control in evolved virtual creatures,
N. Cheney, J. Bongard, V . SunSpiral, and H. Lipson, “On the difficulty of co-optimizing morphology and control in evolved virtual creatures,” in Artificial life conference proceedings . MIT Press One Rogers Street, Cambridge, MA 02142-1209, USA journals-info . . . , 2016, pp. 226–233
2016
-
[19]
Guided self-organization: perception–action loops of embodied systems,
N. Ay, R. Der, and M. Prokopenko, “Guided self-organization: perception–action loops of embodied systems,” Theory in Biosciences, vol. 131, no. 3, pp. 125–127, 2012
2012
-
[20]
Learn goal-conditioned policy with intrinsic motivation for deep reinforcement learning,
J. Liu, D. Wang, Q. Tian, and Z. Chen, “Learn goal-conditioned policy with intrinsic motivation for deep reinforcement learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 36, no. 7, 2022, pp. 7558–7566
2022
-
[21]
Homeokinetic reinforcement learning,
S. C. Smith and J. M. Herrmann, “Homeokinetic reinforcement learning,” in IAPR International Workshop on Partially Supervised Learning. Springer, 2011, pp. 82–91
2011
-
[22]
Evaluation of internal models in autonomous learning,
——, “Evaluation of internal models in autonomous learning,” IEEE Transactions on Cognitive and Developmental Systems , vol. 11, no. 4, pp. 463–472, 2018
2018
-
[23]
Intrinsically motivated goal exploration processes with automatic curriculum learn- ing,
S. Forestier, R. Portelas, Y . Mollard, and P.-Y . Oudeyer, “Intrinsically motivated goal exploration processes with automatic curriculum learn- ing,” Journal of Machine Learning Research , vol. 23, no. 152, pp. 1–41, 2022
2022
-
[24]
Task-agnostic morphology evo- lution,
D. J. Hejna, P. Abbeel, and L. Pinto, “Task-agnostic morphology evo- lution,” in 9th International Conference on Learning Representations, ICLR 2021, 2021
2021
-
[25]
Empowerment: A universal agent-centric measure of control,
A. S. Klyubin, D. Polani, and C. L. Nehaniv, “Empowerment: A universal agent-centric measure of control,” in 2005 ieee congress on evolutionary computation, vol. 1. IEEE, 2005, pp. 128–135
2005
-
[26]
Der and G
R. Der and G. Martius, The playful machine: theoretical foundation and practical realization of self-organizing robots . Springer Science & Business Media, 2012, vol. 15
2012
-
[27]
Improving efficiency of evolving robot designs via self-adaptive learning cycles and an asynchronous archi- tecture,
L. Le Goff and E. Hart, “Improving efficiency of evolving robot designs via self-adaptive learning cycles and an asynchronous archi- tecture,” in Proceedings of the Genetic and Evolutionary Computation Conference Companion, 2024, pp. 1607–1615
2024
-
[28]
Compositional pattern producing networks: A novel abstraction of development,
K. O. Stanley, “Compositional pattern producing networks: A novel abstraction of development,” Genetic programming and evolvable machines, vol. 8, pp. 131–162, 2007
2007
-
[29]
Sample and time efficient policy learning with cma-es and bayesian optimisation,
L. K. Le Goff, E. Buchanan, E. Hart, A. E. Eiben, W. Li, M. De Carlo, M. F. Hale, M. Angus, R. Woolley, J. Timmis, et al. , “Sample and time efficient policy learning with cma-es and bayesian optimisation,” in Artificial Life Conference Proceedings 32 . MIT Press One Rogers St...
2020
-
[30]
Bootstrap- ping artificial evolution to design robots for autonomous fabrication,
E. Buchanan, L. K. Le Goff, W. Li, E. Hart, A. E. Eiben, M. De Carlo, A. F. Winfield, M. F. Hale, R. Woolley, M. Angus, et al., “Bootstrap- ping artificial evolution to design robots for autonomous fabrication,” Robotics, vol. 9, no. 4, p. 106, 2020
2020
-
[31]
Practical hardware for evolvable robots,
M. Angus, E. Buchanan, L. K. Le Goff, E. Hart, A. E. Eiben, M. De Carlo, A. F. Winfield, M. F. Hale, R. Woolley, J. Timmis, et al., “Practical hardware for evolvable robots,” Frontiers in Robotics and AI, vol. 10, p. 1206055, 2023
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.