PF-CD3Q uses online particle filtering to estimate fatigue parameters and constrains a deep Q-learning agent to solve fatigue-aware human-robot task planning as a CMDP.
Acomparativestudyofdeep reinforcement learning models: Dqn vs ppo vs a2c
2 Pith papers cite this work, alongside 12 external citations. Polarity classification is still indexing.
2
Pith papers citing it
12
external citations · external index
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
Applies A2C and PPO agents with post-hoc explanations to optimal battery control in PV-equipped buildings using real LLEC data and shows cost reduction plus policy insights.
citing papers explorer
-
Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production
PF-CD3Q uses online particle filtering to estimate fatigue parameters and constrains a deep Q-learning agent to solve fatigue-aware human-robot task planning as a CMDP.
-
Explainable Data-driven Deep Reinforcement Learning Methods for Optimal Energy Management in Buildings
Applies A2C and PPO agents with post-hoc explanations to optimal battery control in PV-equipped buildings using real LLEC data and shows cost reduction plus policy insights.