REVIEW 1 cited by
Safe Reinforcement Learning From Pixels Using a Stochastic Latent Representation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We address the problem of safe reinforcement learning from pixel observations. Inherent challenges in such settings are (1) a trade-off between reward optimization and adhering to safety constraints, (2) partial observability, and (3) high-dimensional observations. We formalize the problem in a constrained, partially observable Markov decision process framework, where an agent obtains distinct reward and safety signals. To address the curse of dimensionality, we employ a novel safety critic using the stochastic latent actor-critic (SLAC) approach. The latent variable model predicts rewards and safety violations, and we use the safety critic to train safe policies. Using well-known benchmark environments, we demonstrate competitive performance over existing approaches with respects to computational requirements, final reward return, and satisfying the safety constraints.
Forward citations
Cited by 1 Pith paper
-
Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation
Editing a frozen RL policy's latent activations at inference time, using a collision world model, cuts collisions by about 90% on a curated set of hard multirotor scenarios and on real Crazyflies.
Discussion (0). Continue with ORCID to comment.