REVIEW 5 cited by
Phasic Policy Gradient
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce Phasic Policy Gradient (PPG), a reinforcement learning framework which modifies traditional on-policy actor-critic methods by separating policy and value function training into distinct phases. In prior methods, one must choose between using a shared network or separate networks to represent the policy and value function. Using separate networks avoids interference between objectives, while using a shared network allows useful features to be shared. PPG is able to achieve the best of both worlds by splitting optimization into two phases, one that advances training and one that distills features. PPG also enables the value function to be more aggressively optimized with a higher level of sample reuse. Compared to PPO, we find that PPG significantly improves sample efficiency on the challenging Procgen Benchmark.
Forward citations
Cited by 5 Pith papers
-
Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions
In a simulated multi-round ascending takeover auction, toeholds raise the holder's profit but the deterrence channel they are supposed to provide disappears once the contest lasts more than one round.
-
Synthetic emotions and consciousness: exploring architectural boundaries
The paper proposes an architecture that implements emotion-like control without the global broadcast, metarepresentation, autobiographical memory, or cross-module learning that major theories associate with access con...
-
Learning Driven Elastic Task Multi-Connectivity Immersive Computing Systems
A centralized phasic policy gradient agent for elastic VR task offloading with multi-connectivity beats a decentralized independent agent by 28% in latency and 78% in energy in trace-driven simulations.
-
Neural-Enhanced Rate Adaptation and Computation Distribution for Emerging mmWave Multi-User 3D Video Streaming Systems
A cascaded deep RL agent that jointly decides computation placement and bitrate for multi-user mmWave 360 video streaming reports large PSNR and rebuffering gains over fixed-placement ABR baselines in a simulator.
-
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.
Discussion (0). Continue with ORCID to comment.