Pith. sign in

REVIEW 1 cited by

Imitation Learning with Concurrent Actions in 3D Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.05402 v5 pith:MMRBVIC6 submitted 2018-03-14 cs.AI cs.LGstat.ML

classification cs.AIcs.LGstat.ML
keywords learningactionactionsallowscapabilitiescomplexexpertimitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we describe a novel deep reinforcement learning architecture that allows multiple actions to be selected at every time-step in an efficient manner. Multi-action policies allow complex behaviours to be learnt that would otherwise be hard to achieve when using single action selection techniques. We use both imitation learning and temporal difference (TD) reinforcement learning (RL) to provide a 4x improvement in training time and 2.5x improvement in performance over single action selection TD RL. We demonstrate the capabilities of this network using a complex in-house 3D game. Mimicking the behavior of the expert teacher significantly improves world state exploration and allows the agents vision system to be trained more rapidly than TD RL alone. This initial training technique kick-starts TD learning and the agent quickly learns to surpass the capabilities of the expert.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Improved Multi-Agent Algorithm for Cooperative and Competitive Environments by Identifying and Encouraging Cooperation among Agents

    cs.MA 2025-08 reject novelty 4.0 of 10

    Scaling team members' rewards when multiple agents get positive rewards improves MADDPG's empirical team and individual reward in one MPE environment.

Pith tools