Heuresis evaluates six search strategies for autonomous ML research agents and finds that novel ideas are rare, none rated original, and only one reaches top-10 quality while strategies steer axes but do not expand the quality-novelty frontier.
Robots that can adapt like animals
5 Pith papers cite this work, alongside 950 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 5roles
background 2polarities
background 2representative citing papers
EP-based PPO with CPG and residual policies matches standard PPO performance on 12-DoF quadruped uneven-terrain locomotion while using 4.3 times less GPU memory during training.
MONET models multi-task optimization as a task graph and combines neighbor crossover with local mutation, matching or exceeding MAP-Elites baselines on up to 5,000 tasks.
LEAP enables real-time proprioceptive adaptation to unseen damage in a 6DoF soft wrist using HSA actuators by combining latent damage representations with a robust ensemble method, with conditions identified for linear rather than exponential sample complexity.
citing papers explorer
-
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Heuresis evaluates six search strategies for autonomous ML research agents and finds that novel ideas are rare, none rated original, and only one reaches top-10 quality while strategies steer axes but do not expand the quality-novelty frontier.
-
Neuromorphic Reinforcement Learning for Quadruped Locomotion Control on Uneven Terrain
EP-based PPO with CPG and residual policies matches standard PPO performance on 12-DoF quadruped uneven-terrain locomotion while using 4.3 times less GPU memory during training.
-
Multi-Task Optimization over Networks of Tasks
MONET models multi-task optimization as a task graph and combines neighbor crossover with local mutation, matching or exceeding MAP-Elites baselines on up to 5,000 tasks.
-
Damage Adaptation in Seconds for Architected Materials
LEAP enables real-time proprioceptive adaptation to unseen damage in a 6DoF soft wrist using HSA actuators by combining latent damage representations with a robust ensemble method, with conditions identified for linear rather than exponential sample complexity.
- Speculative Rollback Correction for Quality-Diverse Web Agent Imitation