Pith. sign in

arXiv preprint arXiv:2506.14990 , year=

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning, most continual RL papers consider only 3-10 sequential tasks, as CPU-bound environments make longer sequences impractical. Meanwhile, continual learning in cooperative multi-agent settings remains largely unexplored. To address these gaps, we introduce MEAL (Multi-agent Environments for Adaptive Learning), the first benchmark for continual multi-agent RL. By leveraging JAX and GPU acceleration, MEAL enables training on sequences of 100 tasks in a few hours on a single GPU. We find that long task sequences reveal failure modes that do not appear at smaller scales.

years

2026 2 2025 1

verdicts

UNVERDICTED 3

representative citing papers

citing papers explorer

Showing 3 of 3 citing papers.