Pith. sign in

REVIEW 6 cited by

A Survey of Deep Reinforcement Learning in Video Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.10944 v2 pith:PBK46MGZ submitted 2019-12-23 cs.MA cs.AIcs.LG

classification cs.MAcs.AIcs.LG
keywords gameslearningincludingmulti-agentsomevideoachievementsdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep reinforcement learning (DRL) has made great achievements since proposed. Generally, DRL agents receive high-dimensional inputs at each step, and make actions according to deep-neural-network-based policies. This learning mechanism updates the policy to maximize the return with an end-to-end method. In this paper, we survey the progress of DRL methods, including value-based, policy gradient, and model-based algorithms, and compare their main techniques and properties. Besides, DRL plays an important role in game artificial intelligence (AI). We also take a review of the achievements of DRL in various video games, including classical Arcade games, first-person perspective games and multi-agent real-time strategy games, from 2D to 3D, and from single-agent to multi-agent. A large number of video game AIs with DRL have achieved super-human performance, while there are still some challenges in this domain. Therefore, we also discuss some key points when applying DRL methods to this field, including exploration-exploitation, sample efficiency, generalization and transfer, multi-agent learning, imperfect information, and delayed spare rewards, as well as some research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Counterfactual Shapley Credit Assignment

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Counterfactual Shapley values, computed by simulated 'what-if' action replacements, redistribute RL rewards without changing the optimal policy and improve credit assignment in stochastic, sparse, delayed-reward tasks.

  2. Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PEARL trains a policy by differentiating a short horizon of the simulator and using a learned adjoint network to account for long-term return gradients, beating PPO, TD3, BPTT, and SHAC on two double-gyre navigation tasks.

  3. GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning

    cs.LG 2026-01 conditional novelty 6.0 of 10

    GraphAllocBench provides a graph-based resource-allocation benchmark and two supplemental metrics that expose preference-consistency failures in multi-objective RL policies.

  4. Enhancing Reinforcement Learning in 3D Environments through Semantic Segmentation: A Case Study in ViZDoom

    cs.LG 2025-11 conditional novelty 5.0 of 10

    Semantic-segmentation masks can replace RGB input to ViZDoom RL agents with comparable performance and far lower buffer memory, and adding them as an extra channel improves frag scores.

  5. Causal-aware Large Language Models: Enhancing Decision-Making Through Learning, Adapting and Acting

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A learning-adapting-acting framework that uses an LLM to build and update a causal graph of the game world, then uses that graph to generate goals and shape rewards for an RL agent in Crafter.

  6. A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games

    cs.LG 2025-09 conditional novelty 3.0 of 10

    A survey of multi-agent reinforcement learning in video games, plus a proposed five-dimension, MDP-based classification for comparing game complexity.

Pith tools