REVIEW 4 cited by
Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Many real-world tasks involve multiple agents with partial observability and limited communication. Learning is challenging in these settings due to local viewpoints of agents, which perceive the world as non-stationary due to concurrently-exploring teammates. Approaches that learn specialized policies for individual tasks face problems when applied to the real world: not only do agents have to learn and store distinct policies for each task, but in practice identities of tasks are often non-observable, making these approaches inapplicable. This paper formalizes and addresses the problem of multi-task multi-agent reinforcement learning under partial observability. We introduce a decentralized single-task learning approach that is robust to concurrent interactions of teammates, and present an approach for distilling single-task policies into a unified policy that performs well across multiple related tasks, without explicit provision of task identity.
Forward citations
Cited by 4 Pith papers
-
Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning
MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.
-
Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning
Per-agent low-rank adapters on a shared backbone let multi-agent policies specialize at a fraction of the memory cost of separate networks, with competitive benchmark performance.
-
MASK: Multi-Agent Semantic K-Scheduling for Risk-Sensitive 6G Robotics
MASK schedules top-K agents via semantic gating and a global encoder to achieve risk-aware multi-robot coordination that matches unconstrained baselines under bandwidth caps.
-
Orchestrator: Active Inference for Multi-Agent Systems in Long-Horizon Tasks
Orchestrator, an active-inference-inspired feedback system for LLM multi-agent teams, substantially raises maze-solving success rates on medium-difficulty mazes but not consistently on hard mazes.
Discussion (0). Continue with ORCID to comment.