GAMBLe decomposes ADRS into four parameters and an effective landscape, with experiments on 760+ runs across NP-hard problems showing no universal best generator or mechanism and potential gains of 13-67% from component choice.
Title resolution pending
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
EMO-STA evolves a shared program archive across task families then adapts candidates to targets, outperforming matched-compute single-task evolution in most of eight families while reducing overfitting on low-data tasks like ARC.
LIMEN discovers effective RL interfaces by using LLMs to evolve observation and reward programs together from raw state, guided by policy training success, outperforming single-component optimization.
EvoPolicyGym is a new benchmark suite of 16 compact RL environments that evaluates autonomous policy evolution, with GPT-5.5 achieving the top aggregate rank and top-two performance on all tasks.
OPT-BENCH and OPT-Agent evaluate LLM self-optimization in large search spaces, showing stronger models improve via feedback but stay constrained by base capacity and below human performance.
Denoising Recursion Models train multi-step noise reversal in looped transformers and outperform the prior Tiny Recursion Model on ARC-AGI.
citing papers explorer
-
Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
GAMBLe decomposes ADRS into four parameters and an effective landscape, with experiments on 760+ runs across NP-hard problems showing no universal best generator or mechanism and potential gains of 13-67% from component choice.
-
Evolutionary Multi-Task Optimization for LLM-Guided Program Discovery
EMO-STA evolves a shared program archive across task families then adapts candidates to targets, outperforming matched-compute single-task evolution in most of eight families while reducing overfitting on low-data tasks like ARC.
-
Discovering Reinforcement Learning Interfaces with Large Language Models
LIMEN discovers effective RL interfaces by using LLMs to evolve observation and reward programs together from raw state, guided by policy training success, outperforming single-component optimization.
-
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
EvoPolicyGym is a new benchmark suite of 16 compact RL environments that evaluates autonomous policy evolution, with GPT-5.5 achieving the top aggregate rank and top-two performance on all tasks.
-
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
OPT-BENCH and OPT-Agent evaluate LLM self-optimization in large search spaces, showing stronger models improve via feedback but stay constrained by base capacity and below human performance.
-
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
Denoising Recursion Models train multi-step noise reversal in looped transformers and outperform the prior Tiny Recursion Model on ARC-AGI.