Pith. sign in

REVIEW 6 cited by

Human-Timescale Adaptation in an Open-Ended Task Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.07608 v1 pith:DO3MJ5XE submitted 2023-01-18 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords learningadaptationagentopen-endedtaskacrossadaptivedemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models have shown impressive adaptation and scalability in supervised and self-supervised learning problems, but so far these successes have not fully translated to reinforcement learning (RL). In this work, we demonstrate that training an RL agent at scale leads to a general in-context learning algorithm that can adapt to open-ended novel embodied 3D problems as quickly as humans. In a vast space of held-out environment dynamics, our adaptive agent (AdA) displays on-the-fly hypothesis-driven exploration, efficient exploitation of acquired knowledge, and can successfully be prompted with first-person demonstrations. Adaptation emerges from three ingredients: (1) meta-reinforcement learning across a vast, smooth and diverse task distribution, (2) a policy parameterised as a large-scale attention-based memory architecture, and (3) an effective automated curriculum that prioritises tasks at the frontier of an agent's capabilities. We demonstrate characteristic scaling laws with respect to network size, memory length, and richness of the training task distribution. We believe our results lay the foundation for increasingly general and adaptive RL agents that perform well across ever-larger open-ended domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 22 citations worldwide. Full citation record

  1. RoboTTT: Context Scaling for Robot Policies

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A robot policy that updates its own weights during deployment can use 8,000 steps of history, steadily improving as context grows and enabling one-shot imitation from human videos.

  2. Cross-Entropy Games for Language Models: From Implicit Knowledge to General Capability Measures

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Xent Games formalize a large family of LLM evaluation tasks as games whose rewards and constraints are signed cross-entropy sums, and propose using them to build general capability measures.

  3. HERAKLES: Hierarchical Skill Compilation for Open-ended LLM Agents

    cs.LG 2025-08 conditional novelty 6.0 of 10

    HERAKLES couples a language-model planner to a small, continually retrained skill executor and outperforms three baselines on the 17-goal Crafter benchmark, scaling better to reworded and repeated goals.

  4. Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments

    cs.LG 2026-03 conditional novelty 5.0 of 10

    PPO plateaus can be avoided by increasing the number of parallel environments, which reduces both the outer-loop step size and update noise; scaling to 1M environments sustained improvement to 1T transitions.

  5. Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Recurrent neural policies trained on episodic tasks converge to stable cyclic attractors in hidden state, and the geometry of these cycles mirrors behavior structure.

  6. Training Cross-Morphology Embodied AI Agents: From Practical Challenges to Theoretical Foundations

    cs.AI 2025-06 conditional novelty 3.0 of 10

    The paper proves that the cross-morphology robot training problem HEAT is PSPACE-complete by embedding any POMDP into a HEAT instance with a single morphology.

Pith tools