Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:31.896668Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2501.12627.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:31.896668Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T12:15:08.304150Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.723680Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bd12ae0c-9e7c-47c1-ba72-949a942c50d9 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Deep reinforcement learning at the edge of the statistical precipice
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 777d5ba1-acf0-467c-9d2e-420e08bb6f61 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Atari-5: Distilling the arcade learning environment down to five games
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b025b990-4e9d-4136-91ad-7e30cf51b870 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Existence, relatedness, and growth: Human needs in organizational settings
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fab2e43a-96cf-48ae-900d-3b37447c7a19 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Using confidence bounds for exploitation-exploration trade-offs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 42f52b22-413a-4239-ac7e-ae3ce297336c · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Never give up: Learning directed exploration strategies
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc46cc3b-895d-4b3d-8121-fd8c24dbfd99 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model The arcade learning environment: An evaluation platform for general agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1553f4ef-748c-4999-a2d7-b454bfa9b33e · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Unifying count-based exploration and intrinsic motivation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29042ab4-e56a-47ea-a5dc-495eafe0d9fc · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A markovian decision process
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d1afa23-e649-41e7-83cf-a20d2800927d · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration by random network distillation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80fddff9-f795-4460-b351-96a5501b9a18 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Explore, discover and learn: Unsupervised discovery of state-covering skills
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10b4dfd5-e1dd-453f-95c2-e8e02292dbfc · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85ba021e-a615-4c9e-9059-04ffd982424a · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Leveraging procedural generation to benchmark reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b48be9ad-2e89-446a-857f-a7f686b64864 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Stochastic linear optimization under bandit feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0dabe982-1812-4302-9b1a-383c4fc6b0af · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Diversity is all you need: Learning skills without a reward function
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d996408c-e6e1-41f5-9480-c491fba766b4 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Adversarially guided actor-critic
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c02a526d-a851-4081-a54a-850e1efcf554 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Variational Intrinsic Control
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f8d2c1-69c4-40ef-b290-4a966b455c1f · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Fast task inference with variational intrinsic successor features
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40c39387-4a1a-4ff8-a700-7ce1765cfaae · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Provably efficient maximum entropy exploration
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c77159e-7fab-45ef-a465-6146559a7f82 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration via elliptical episodic bonuses
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecdcb10c-28dd-4b67-8e33-85840a9f89e1 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 726615b0-541b-47b7-85c3-e284cc80958b · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Planning and acting in partially observable stochastic domains
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb2441da-0e0b-4c45-b9b7-92c3c81f1039 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curl: Contrastive unsupervised representations for reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b80cb0b-265f-46b1-a407-77add88a25d5 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Cic: Contrastive intrinsic control for unsupervised skill discovery
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b57d581-455a-4b5d-9b79-286e760ea403 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Urlb: Unsupervised reinforcement learning benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbfafc83-6a9c-4275-aec1-51134eeed3cd · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A contextual-bandit approach to personalized news article recommendation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586a0342-e2ce-4c97-a19a-f321767e38ac · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Aps: Active pretraining with successor features
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9c13c34-23d2-40fa-9e4e-a9a9e1d6d75e · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with the successor representation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 167bf6db-1c33-49a3-a2db-506bafe739d2 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A dynamic theory of human motivation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e37be31-db05-4499-93aa-fec1ee4c345b · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Improving intrinsic exploration with language abstractions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0e6ea44-e187-4119-b82d-106f8e2f4acb · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with neural density models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 486ab38a-5025-4edb-b7e2-680f42fbb893 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Lipschitz-constrained Unsupervised Skill Discovery
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d051dbc2-2728-4121-bad0-cf8043d7f72c · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curiosity-driven exploration by self-supervised prediction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae2e2b5-728e-483f-a7c1-cdf48f88d5ec · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Self-supervised exploration via disagreement
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c253f65d-1884-4552-b4aa-34d7aeb1d127 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Ride: Rewarding impact-driven exploration for procedurally-generated environments
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e7f619b6-92d2-4171-836b-d697f8a232cb · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minihack the planet: A sandbox for open-ended reinforcement learning research
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51ccc0d7-1fd0-4e2b-b73b-846d65d96fa9 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e019307f-01e9-42f1-9754-8ce787f3b7cb · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model State entropy maximization with random encoders for efficient exploration
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bd73490-d8c9-4f91-baf4-0b9fe55ed8fc · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c5f284-6475-4019-b3f9-cefc29542115 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning: An introduction
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cef02a3-edf5-43c6-9fc8-4a27ada1fc55 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model \# exploration: A study of count-based exploration for deep reinforcement learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bcbb78-6710-494f-8f9a-a84688d5851a · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning with prototypical representations
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce6a0bb6-ad98-4bdf-8008-b911a70e1984 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rewarding episodic visitation discrepancy for exploration in reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f4e8c4b-f613-45e3-8860-13a13a60dd75 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model R \'e nyi state entropy maximization for exploration acceleration in reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3e1a87a-8f8a-406a-abc1-4777764238e2 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892dd3d4-1673-431a-b114-762c796b0303 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rllte: Long-term evolution project of reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9f54259-2184-4fac-8134-bcbf9206621d · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Noveld: A simple yet effective exploration criterion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a0e2ef72-6f7e-4268-8b52-866ea9ee5639 · outbound
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model write newline
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea666f1f-983f-42cd-9966-94b2d78a1c1d · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Deep Reinforcement Learning with Hybrid Intrinsic Reward Model
Reference 251
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.