LPM uses a dual-network design to compute intrinsic rewards from the change in prediction error across iterations, providing a noise-robust signal that is theoretically linked to information gain.
Bayesian curiosity for efficient exploration in reinforce- ment learning.arXiv preprint arXiv:1911.08701
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
A revised DMBN with positional time encoding improves temporal representation and generalization in neural processes for multimodal robotic action prediction.
citing papers explorer
-
Beyond Noisy-TVs: Noise-Robust Exploration Via Learning Progress Monitoring
LPM uses a dual-network design to compute intrinsic rewards from the change in prediction error across iterations, providing a noise-robust signal that is theoretically linked to information gain.
-
Exploring Temporal Representation in Neural Processes for Multimodal Action Prediction
A revised DMBN with positional time encoding improves temporal representation and generalization in neural processes for multimodal robotic action prediction.