Pith. sign in

REVIEW 1 cited by

Learning to Cooperate with Unseen Agent via Meta-Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.03431 v1 pith:XPOR4RDW submitted 2021-11-05 cs.AI cs.LGcs.MA

classification cs.AIcs.LGcs.MA
keywords cooperativeagentlearningagentscooperatedomainknowledgemeta-reinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ad hoc teamwork problem describes situations where an agent has to cooperate with previously unseen agents to achieve a common goal. For an agent to be successful in these scenarios, it has to have a suitable cooperative skill. One could implement cooperative skills into an agent by using domain knowledge to design the agent's behavior. However, in complex domains, domain knowledge might not be available. Therefore, it is worthwhile to explore how to directly learn cooperative skills from data. In this work, we apply meta-reinforcement learning (meta-RL) formulation in the context of the ad hoc teamwork problem. Our empirical results show that such a method could produce robust cooperative agents in two cooperative environments with different cooperative circumstances: social compliance and language interpretation. (This is a full paper of the extended abstract version.)

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of In-Context Reinforcement Learning

    cs.LG 2025-02 unverdicted novelty 3.0 of 10

    Recent in-context reinforcement learning literature is organized into a taxonomy covering pretraining methods, context construction, test-time performance, theory, and architectures, with no new experiments or theory ...

Pith tools