REVIEW 12 cited by
OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Open-ended and AI-generating algorithms aim to continuously generate and solve increasingly complex tasks indefinitely, offering a promising path toward more general intelligence. To accomplish this grand vision, learning must occur within a vast array of potential tasks. Existing approaches to automatically generating environments are constrained within manually predefined, often narrow distributions of environment, limiting their ability to create any learning environment. To address this limitation, we introduce a novel framework, OMNI-EPIC, that augments previous work in Open-endedness via Models of human Notions of Interestingness (OMNI) with Environments Programmed in Code (EPIC). OMNI-EPIC leverages foundation models to autonomously generate code specifying the next learnable (i.e., not too easy or difficult for the agent's current skill set) and interesting (e.g., worthwhile and novel) tasks. OMNI-EPIC generates both environments (e.g., an obstacle course) and reward functions (e.g., progress through the obstacle course quickly without touching red objects), enabling it, in principle, to create any simulatable learning task. We showcase the explosive creativity of OMNI-EPIC, which continuously innovates to suggest new, interesting learning challenges. We also highlight how OMNI-EPIC can adapt to reinforcement learning agents' learning progress, generating tasks that are of suitable difficulty. Overall, OMNI-EPIC can endlessly create learnable and interesting environments, further propelling the development of self-improving AI systems and AI-Generating Algorithms. Project website with videos: https://dub.sh/omniepic
Forward citations
Cited by 12 Pith papers
-
Octax: Accelerated CHIP-8 Arcade Environments for Reinforcement Learning in JAX
Octax is a JAX-based CHIP-8 emulator that runs thousands of parallel arcade environments on GPUs (350k steps/s) and supports LLM-generated games for RL training.
-
How Should We Meta-Learn Reinforcement Learning Algorithms?
A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...
-
To Trade or Not to Trade: An Agentic Approach to Estimating Market Risk Improves Trading Decisions
LLM-discovered stochastic models of price paths provide risk metrics that improve trader-agent decisions, raising average Sharpe ratios from 0.88 to 1.40 in the paper's backtests.
-
Self-Challenging Language Model Agents
A language model agent can generate its own verifiable training tasks and improve its tool-use success rate by about 2x without human-annotated data.
-
Exploring Flow-Lenia Universes with a Curiosity-driven AI Scientist: Discovering Diverse Ecosystem Dynamics
IMGEP goal exploration on simulation-wide metrics discovers more diverse Flow-Lenia ecosystem and matter-movement dynamics than random search, though key metric details and a claimed scaling study are missing.
-
Automated Capability Discovery via Foundation Model Self-Exploration
ACD automatically generates thousands of open-ended tasks and clusters them into dozens of capability and failure categories, with LLM-vs-human scoring agreement (F1 = 0.86).
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
The paper proposes that self-improving agents must learn to manage their own learning processes, framing this as intrinsic metacognitive learning, and argues it is necessary for sustained and generalized improvement.
-
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
EXIF repeatedly has a teacher agent explore an environment, relabel the exploration as tasks, train a student agent on it, and use the student's failures to guide the next round, improving 7B-8B agents in Webshop and Crafter.
-
Towards a Formal Theory of the Need for Competence via Computational Intrinsic Motivation
Four facets of competence in Self-Determination Theory can be matched to existing reinforcement learning formalisms, revealing hidden assumptions in the psychological theory.
-
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games
An LM-driven feedback loop that tunes reward weights from scalar performance statistics reaches 80.4% lap success in a racing task, close to a human expert's peak of 93.6%.
-
VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
VoyagerVision combines GPT-4o with point-of-view screenshots in the Voyager Minecraft agent, producing 18 verified unit-test building successes out of 50 attempts and 2.75 unique structures per 50-step open-ended run.
Discussion (0). Continue with ORCID to comment.