Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.
A survey of generalisation in deep reinforcement learning
3 Pith papers cite this work, alongside 13 external citations. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
MAPLE proposes latent multi-agent rollouts with supervised fine-tuning followed by reinforcement learning using safety, progress, interaction, and diversity rewards to enable scalable closed-loop training for end-to-end autonomous driving.
XIPER creates a reward signal for cross-domain video imitation learning by training a video prediction model that maps agent views to the expert domain and scoring prediction likelihood.
citing papers explorer
-
Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.
-
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
MAPLE proposes latent multi-agent rollouts with supervised fine-tuning followed by reinforcement learning using safety, progress, interaction, and diversity rewards to enable scalable closed-loop training for end-to-end autonomous driving.
-
Reinforcement Learning from Cross-domain Videos with Video Prediction Model
XIPER creates a reward signal for cross-domain video imitation learning by training a video prediction model that maps agent views to the expert domain and scoring prediction likelihood.