Pith. sign in

REVIEW 1 cited by

Emergence of Communication in an Interactive World with Consistent Speakers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1809.00549 v2 pith:FRKJOIBO submitted 2018-09-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords traininggradientpolicycommunicationagentsalgorithmcomparedconsistent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training agents to communicate with one another given task-based supervision only has attracted considerable attention recently, due to the growing interest in developing models for human-agent interaction. Prior work on the topic focused on simple environments, where training using policy gradient was feasible despite the non-stationarity of the agents during training. In this paper, we present a more challenging environment for testing the emergence of communication from raw pixels, where training using policy gradient fails. We propose a new model and training algorithm, that utilizes the structure of a learned representation space to produce more consistent speakers at the initial phases of training, which stabilizes learning. We empirically show that our algorithm substantially improves performance compared to policy gradient. We also propose a new alignment-based metric for measuring context-independence in emerged communication and find our method increases context-independence compared to policy gradient and other competitive baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In Pursuit of Predictive Models of Human Preferences Toward AI Teammates

    cs.HC 2025-01 reject novelty 5.0 of 10

    In a 241-participant Hanabi study, AI behavioral metrics like action diversity and strategic dominance predict human preference ratings more strongly than the final game score, though all correlations are weak to moderate.

Pith tools