REVIEW 9 cited by
Noisy Networks for Exploration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet is straightforward to implement and adds little computational overhead. We find that replacing the conventional exploration heuristics for A3C, DQN and dueling agents (entropy reward and $\epsilon$-greedy respectively) with NoisyNet yields substantially higher scores for a wide range of Atari games, in some cases advancing the agent from sub to super-human performance.
Forward citations
Cited by 9 Pith papers
-
Prompt-Driven Exploration
Prompt-Driven Exploration refines language prompts from rollout videos via a VLM, enabling RL to escape zero-reward VLA and LLM policies where action-space noise fails.
-
Skillful joint probabilistic weather forecasting from marginals
FGN, a neural weather model trained only on per-location forecast scores, produces more accurate global ensemble forecasts than GenCast and captures realistic spatial correlations.
-
How Should We Meta-Learn Reinforcement Learning Algorithms?
A systematic comparison of black-box evolution, neural and symbolic distillation, and LLM-based proposal for meta-learning RL algorithms yields practical recommendations: warm-started LLM proposal is sample-efficient,...
-
Meta-learning how to Share Credit among Macro-Actions
MASP meta-learns a similarity matrix over macro-actions and regularizes Q-values so that similar actions move together, improving exploration and performance in augmented-action-space RL.
-
When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design
Dueling DQN is the only tested value-based RL variant that reliably improves seven-variable PCSEL designs under a matched 83-call FDTD budget across four seeds.
-
A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
Tunable energy landscapes whose thermal averages equal sigmoid, softmax, and matrix-vector products can, in principle, form the basis of a low-energy analog computer, with a superconducting double-well device as a fir...
-
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
Bayesian deep Q-learning exhibits a cold posterior effect, caused partly by misspecified Gaussian priors, and Laplace or meta-learned priors improve performance.
-
Quantum Reinforcement Learning by Adaptive Non-local Observables
Adaptive non-local observables, jointly trained with variational circuit parameters, improve DQN and A3C reinforcement learning agents on simulated benchmark tasks relative to fixed Pauli-measurement baselines.
-
Designing Adaptive Algorithms Based on Reinforcement Learning for Dynamic Optimization of Sliding Window Size in Multi-Dimensional Data Streams
A DQN-based reinforcement learning agent that dynamically chooses sliding window sizes is claimed to improve classification accuracy on multi-dimensional streams, but the reported evaluation is inconsistent and not re...
Discussion (0). Sign in to comment.