Closed-loop prompt-based translation with hierarchical verification and iterative repair produces equivalent high-performance RL environments across five cases including new TCGJax.
Amago: Scalable in-context reinforcement learning for adaptive agents.arXiv preprint arXiv:2310.09971,
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 3years
2026 3representative citing papers
A multi-task multi-modal transformer policy pretrained on offline trajectories from thousands of RL environments matches task-specific reference policies on approximately 1000 environments spanning robotics, driving, inventory, cybersecurity, trading, and games.
A Graph Attention Network pretrained solely on synthetic MDPs solves held-out tabular RL benchmarks in context, outperforming UCB-VI and Q-learning online while matching VI-LCB offline.
citing papers explorer
-
Automatic Generation of High-Performance RL Environments
Closed-loop prompt-based translation with hierarchical verification and iterative repair produces equivalent high-performance RL environments across five cases including new TCGJax.
-
Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models
A multi-task multi-modal transformer policy pretrained on offline trajectories from thousands of RL environments matches task-specific reference policies on approximately 1000 environments spanning robotics, driving, inventory, cybersecurity, trading, and games.
-
Reinforcement Learning Foundation Models Should Already Be A Thing
A Graph Attention Network pretrained solely on synthetic MDPs solves held-out tabular RL benchmarks in context, outperforming UCB-VI and Q-learning online while matching VI-LCB offline.