Pith. sign in

REVIEW 2 cited by

Discriminative Particle Filter Reinforcement Learning for Complex Partial Observations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.09884 v1 pith:HBRPKQWO submitted 2020-02-23 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords observationscomplexparticledecisiondiscriminativedpfrlfilterlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep reinforcement learning is successful in decision making for sophisticated games, such as Atari, Go, etc. However, real-world decision making often requires reasoning with partial information extracted from complex visual observations. This paper presents Discriminative Particle Filter Reinforcement Learning (DPFRL), a new reinforcement learning framework for complex partial observations. DPFRL encodes a differentiable particle filter in the neural network policy for explicit reasoning with partial observations over time. The particle filter maintains a belief using learned discriminative update, which is trained end-to-end for decision making. We show that using the discriminative update instead of standard generative models results in significantly improved performance, especially for tasks with complex visual observations, because they circumvent the difficulty of modeling complex observations that are irrelevant to decision making. In addition, to extract features from the particle belief, we propose a new type of belief feature based on the moment generating function. DPFRL outperforms state-of-the-art POMDP RL models in Flickering Atari Games, an existing POMDP RL benchmark, and in Natural Flickering Atari Games, a new, more challenging POMDP RL benchmark introduced in this paper. Further, DPFRL performs well for visual navigation with real-world data in the Habitat environment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Model-Based Reinforcement Learning under Random Observation Delays

    cs.LG 2025-09 unverdicted novelty 6.0 of 10

    A delay-aware model-based RL framework with sequential belief filtering handles random out-of-sequence observations in POMDPs and outperforms MDP baselines while showing robustness to delay shifts.

  2. Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A PPO policy selecting one sensor per second from particle-filter belief features is statistically equivalent to explicit information-gain selection and non-inferior to all-sensors-on within 2 m RMSE / 2% lost-track m...

Pith tools