DreamerV2 reaches human-level performance on 55 Atari games by learning behaviors inside a separately trained discrete-latent world model.
Noisy networks for exploration
6 Pith papers cite this work, alongside 390 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Adversarial RL approximates a game-theoretic equilibrium to yield a stochastic policy for prioritizing alerts against adaptive attackers in fraud and intrusion detection.
LEDE reframes speculative decoding as an MDP and applies offline RL to learn dynamic policies for exit layer and speculation length selection, delivering 2.0-2.7x speedups over autoregressive decoding on Llama-2/3 models.
Evolvability ES is an evolutionary strategy variant that directly optimizes for evolvability by maximizing behavioral diversity under mutations, tested on 2D/3D locomotion tasks and shown competitive with MAML.
PF-CD3Q uses online particle filtering to estimate fatigue parameters and constrains a deep Q-learning agent to solve fatigue-aware human-robot task planning as a CMDP.
Rainbow DQN with kinematics-aware design optimization enables reliable cooperative insertion by Delta and 3-RRS robots in a high-fidelity simulator.
citing papers explorer
-
Mastering Atari with Discrete World Models
DreamerV2 reaches human-level performance on 55 Atari games by learning behaviors inside a separately trained discrete-latent world model.
-
Finding Needles in a Moving Haystack: Prioritizing Alerts with Adversarial Reinforcement Learning
Adversarial RL approximates a game-theoretic equilibrium to yield a stochastic policy for prioritizing alerts against adaptive attackers in fraud and intrusion detection.
-
Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning
LEDE reframes speculative decoding as an MDP and applies offline RL to learn dynamic policies for exit layer and speculation length selection, delivering 2.0-2.7x speedups over autoregressive decoding on Llama-2/3 models.
-
Evolvability ES: Scalable and Direct Optimization of Evolvability
Evolvability ES is an evolutionary strategy variant that directly optimizes for evolvability by maximizing behavioral diversity under mutations, tested on 2D/3D locomotion tasks and shown competitive with MAML.
-
Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production
PF-CD3Q uses online particle filtering to estimate fatigue parameters and constrains a deep Q-learning agent to solve fatigue-aware human-robot task planning as a CMDP.
-
Rainbow Deep Q-Learning with Kinematics-Aware Design for Cooperative Delta and 3-RRS Parallel Robot Insertion
Rainbow DQN with kinematics-aware design optimization enables reliable cooperative insertion by Delta and 3-RRS robots in a high-fidelity simulator.