Pith. sign in

REVIEW 2 cited by

MAMBA: an Effective World Model Approach for Meta-Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09859 v1 pith:K5MGXJY4 submitted 2024-03-14 cs.LG

classification cs.LG
keywords meta-rlapproachdomainsmodel-basedchallengingefficiencyexistinglearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Meta-reinforcement learning (meta-RL) is a promising framework for tackling challenging domains requiring efficient exploration. Existing meta-RL algorithms are characterized by low sample efficiency, and mostly focus on low-dimensional task distributions. In parallel, model-based RL methods have been successful in solving partially observable MDPs, of which meta-RL is a special case. In this work, we leverage this success and propose a new model-based approach to meta-RL, based on elements from existing state-of-the-art model-based and meta-RL methods. We demonstrate the effectiveness of our approach on common meta-RL benchmark domains, attaining greater return with better sample efficiency (up to $15\times$) while requiring very little hyperparameter tuning. In addition, we validate our approach on a slate of more challenging, higher-dimensional domains, taking a step towards real-world generalizing agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Success in Humanoid Reinforcement Learning under Partial Observation

    cs.AI 2025-07 reject novelty 6.0 of 10

    This paper reports the first stable partial-observability training on Humanoid-v4, using a parallel history encoder that matches full-state TD3 performance in most tested state-removal settings.

  2. Embodied Intelligence: The Key to Unblocking Generalized Artificial Intelligence

    cs.AI 2025-05 conditional novelty 2.0 of 10

    This review argues that embodied intelligence is the essential route to AGI and analyzes four modular components, but it adds no new experimental or theoretical results.

Pith tools