Pith. sign in

REVIEW 2 cited by

A Comparative Study of Deep Reinforcement Learning Models: DQN vs PPO vs A2C

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14151 v1 pith:77MB5KML submitted 2024-07-19 cs.LG

classification cs.LG
keywords learningmodelscomparativedeepgamereinforcementactor-criticadaptability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study conducts a comparative analysis of three advanced Deep Reinforcement Learning models: Deep Q-Networks (DQN), Proximal Policy Optimization (PPO), and Advantage Actor-Critic (A2C), within the BreakOut Atari game environment. Our research assesses the performance and effectiveness of these models in a controlled setting. Through rigorous experimentation, we examine each model's learning efficiency, strategy development, and adaptability under dynamic game conditions. The findings provide critical insights into the practical applications of these models in game-based learning environments and contribute to the broader understanding of their capabilities. The code is publicly available at github.com/Neilus03/DRL_comparative_study.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    PF-CD3Q uses online particle filtering to estimate fatigue parameters and constrains a deep Q-learning agent to solve fatigue-aware human-robot task planning as a CMDP.

  2. Explainable Data-driven Deep Reinforcement Learning Methods for Optimal Energy Management in Buildings

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    Applies A2C and PPO agents with post-hoc explanations to optimal battery control in PV-equipped buildings using real LLEC data and shows cost reduction plus policy insights.

Pith tools