Pith. sign in

Paper Citation Record · LEDGER

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2501.12627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12627 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:04:31.896668Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T12:15:08.304150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.723680Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd12ae0c-9e7c-47c1-ba72-949a942c50d9 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.671849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.664680Z digest=sha256:cedc67c6a3d4dafb06105c296d5fd4305fb13d919866978ae9e1c9622bd61c42

Observation 777d5ba1-acf0-467c-9d2e-420e08bb6f61 · outbound

This paper cites Atari-5: Distilling the arcade learning environment down to five games.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Atari-5: Distilling the arcade learning environment down to five games

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.655050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.670067Z digest=sha256:ee910cb18c92a1289804fb3947b3af16cfea1ff18bd67427586231130d926ae7

Observation b025b990-4e9d-4136-91ad-7e30cf51b870 · outbound

This paper cites Existence, relatedness, and growth: Human needs in organizational settings.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Existence, relatedness, and growth: Human needs in organizational settings

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.639893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.674886Z digest=sha256:a5d3a7093f15d202b93eac5c9172d6149a0426dd39f5cbff6419af403d9bc20f

Observation fab2e43a-96cf-48ae-900d-3b37447c7a19 · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Using confidence bounds for exploitation-exploration trade-offs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.624106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.679787Z digest=sha256:1a1260ed50b9df5742b1e462b255991e9fe0b5703b01e6cd7d93809eadac098e

Observation 42f52b22-413a-4239-ac7e-ae3ce297336c · outbound

This paper cites Never give up: Learning directed exploration strategies.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Never give up: Learning directed exploration strategies

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.608864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.684579Z digest=sha256:56df80003dd315131b07ed985615cc33c9ed4ea0558d68562b440d0a4aecc5db

Observation dc46cc3b-895d-4b3d-8121-fd8c24dbfd99 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model The arcade learning environment: An evaluation platform for general agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.593250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.689560Z digest=sha256:de7aba21b2a462544b4366beaf311096a3808e8eb73fb999e83f97232a6edce3

Observation 1553f4ef-748c-4999-a2d7-b454bfa9b33e · outbound

This paper cites Unifying count-based exploration and intrinsic motivation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Unifying count-based exploration and intrinsic motivation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.578496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.694801Z digest=sha256:41901426bb00c9090ef4494d0d41253059298115c2aa90b1865c49b8e91d2505

Observation 29042ab4-e56a-47ea-a5dc-495eafe0d9fc · outbound

This paper cites A markovian decision process.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A markovian decision process

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.562582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.699397Z digest=sha256:8299faa5684ffbc8a4fa8f6d0692a7eecf3e112b23a099be1a255847dff6cc44

Observation 9d1afa23-e649-41e7-83cf-a20d2800927d · outbound

This paper cites Exploration by random network distillation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration by random network distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.546712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.705179Z digest=sha256:123d30ab58249381af473cbb085ac2f8e56a96c49783606d569cb261967886ef

Observation 80fddff9-f795-4460-b351-96a5501b9a18 · outbound

This paper cites Explore, discover and learn: Unsupervised discovery of state-covering skills.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Explore, discover and learn: Unsupervised discovery of state-covering skills

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.531059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.710552Z digest=sha256:33e3f0cc2f6bf299481d875f064742def1cc3e1b2f53f9a154c21756a0a601fa

Observation 10b4dfd5-e1dd-453f-95c2-e8e02292dbfc · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.514663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.715540Z digest=sha256:831cbdbccb53659b113de79681cea567aee2162ac6c9532dfa1e952093839115

Observation 85ba021e-a615-4c9e-9059-04ffd982424a · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Leveraging procedural generation to benchmark reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.497491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.720712Z digest=sha256:5cd9677aba0622d19f60fc85a938f7677f5dd1ee29ac3cbe1c2b53e7d2f41b4b

Observation b48be9ad-2e89-446a-857f-a7f686b64864 · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Stochastic linear optimization under bandit feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.479715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.725530Z digest=sha256:545a11c874de905c1b299bfb8744c9d2601b05d0780e5ce156dcf4d212bdc92d

Observation 0dabe982-1812-4302-9b1a-383c4fc6b0af · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Diversity is all you need: Learning skills without a reward function

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.464273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.730698Z digest=sha256:027d1e829d486dfd5f15d7ce89d4f7af822db15a45942eab60d7009e0f67cfed

Observation d996408c-e6e1-41f5-9480-c491fba766b4 · outbound

This paper cites Adversarially guided actor-critic.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Adversarially guided actor-critic

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.448723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.735368Z digest=sha256:5c18e0f82aea08f9ea4d1424ce9370854a120bb4f2e206cedf3ee8f862e42d8f

Observation c02a526d-a851-4081-a54a-850e1efcf554 · outbound

This paper cites Variational Intrinsic Control.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Variational Intrinsic Control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.740034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.740034Z digest=sha256:94efcb574b690df94bc1944ac0520b82618d3f696207620ee71c8c8e2395c61a

Observation 02f8d2c1-69c4-40ef-b290-4a966b455c1f · outbound

This paper cites Fast task inference with variational intrinsic successor features.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Fast task inference with variational intrinsic successor features

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.432093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.745680Z digest=sha256:1e18f030d1f013b09533243c1e2182f63d7b9918e7dc158f4bc066fc1d02465b

Observation 40c39387-4a1a-4ff8-a700-7ce1765cfaae · outbound

This paper cites Provably efficient maximum entropy exploration.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Provably efficient maximum entropy exploration

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.414030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.750439Z digest=sha256:8cbf05c29dde33eff4148b18bd9832ea9c4ad1c2bcb6040846b626f9e73da77a

Observation 6c77159e-7fab-45ef-a465-6146559a7f82 · outbound

This paper cites Exploration via elliptical episodic bonuses.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Exploration via elliptical episodic bonuses

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.396886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.755324Z digest=sha256:f6f263c24a40e6c02eb08f8f78cf53ac4da8eaba3e3de076716d976f71cf6289

Observation ecdcb10c-28dd-4b67-8e33-85840a9f89e1 · outbound

This paper cites A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:04:32.011275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.760641Z digest=sha256:65b375790b16f1ff5fe71e7e6995276af67ad060fd261b32bcde30431d0ca800

Observation 726615b0-541b-47b7-85c3-e284cc80958b · outbound

This paper cites Planning and acting in partially observable stochastic domains.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Planning and acting in partially observable stochastic domains

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.380601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.765980Z digest=sha256:3f2c6bd9af715a43b8d2b3c84f3f1958c721456b728f80fab9079c1d9791bd95

Observation bb2441da-0e0b-4c45-b9b7-92c3c81f1039 · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curl: Contrastive unsupervised representations for reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.365128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.770736Z digest=sha256:e8b5fc2a6de31bc48cea56607113c7c47e9c052f2a5c0a5460c64c3842cda6f1

Observation 2b80cb0b-265f-46b1-a407-77add88a25d5 · outbound

This paper cites Cic: Contrastive intrinsic control for unsupervised skill discovery.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Cic: Contrastive intrinsic control for unsupervised skill discovery

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.349486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.775401Z digest=sha256:9b4799fa76e4681668b6e7418033dea2861c799fd9dac2e6f9ef3e5a455a4cb1

Observation 0b57d581-455a-4b5d-9b79-286e760ea403 · outbound

This paper cites Urlb: Unsupervised reinforcement learning benchmark.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Urlb: Unsupervised reinforcement learning benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.333107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.780024Z digest=sha256:7529348f3edaff0daf4a166944dc99ea79df5e94471cc8ceea92d178e3fc42fe

Observation fbfafc83-6a9c-4275-aec1-51134eeed3cd · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A contextual-bandit approach to personalized news article recommendation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.785023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.785023Z digest=sha256:76844bb24a9312ae479c0a2b53f8a94dcf7abd41ae688d445a65dfe7b391bb60

Observation 586a0342-e2ce-4c97-a19a-f321767e38ac · outbound

This paper cites Aps: Active pretraining with successor features.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Aps: Active pretraining with successor features

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.305952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.789842Z digest=sha256:92e38da449d13b42f3752c2f7e516f0e0e02df8833cde267ac22fe5535d570df

Observation f9c13c34-23d2-40fa-9e4e-a9a9e1d6d75e · outbound

This paper cites Count-based exploration with the successor representation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with the successor representation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.290227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.794533Z digest=sha256:15d018a44b2e664043b64f61d3b44eb7db664d0b70cd1bcf99c75b7f895a836c

Observation 167bf6db-1c33-49a3-a2db-506bafe739d2 · outbound

This paper cites A dynamic theory of human motivation.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model A dynamic theory of human motivation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.274824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.801075Z digest=sha256:f3ce7b92572fc2ac3c244fc9afb802f5879b52b44047a2d61448f0736fabbdaf

Observation 3e37be31-db05-4499-93aa-fec1ee4c345b · outbound

This paper cites Improving intrinsic exploration with language abstractions.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Improving intrinsic exploration with language abstractions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.258772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.806136Z digest=sha256:b8b516f1f18702e1ec9493ba2527b5a74fb5399cde546a5a93f46792bcc5e839

Observation f0e6ea44-e187-4119-b82d-106f8e2f4acb · outbound

This paper cites Count-based exploration with neural density models.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Count-based exploration with neural density models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.242416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.811005Z digest=sha256:08b13672b9f4d75ae29dd294eaa724dbf13c4a2858565ec8acb2219368539cbf

Observation 486ab38a-5025-4edb-b7e2-680f42fbb893 · outbound

This paper cites Lipschitz-constrained Unsupervised Skill Discovery.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Lipschitz-constrained Unsupervised Skill Discovery

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.816155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.816155Z digest=sha256:64f710733bdd1e05f5ae307e4bd133c9687af7114676a7d46c6d9a7dc7fe60d0

Observation d051dbc2-2728-4121-bad0-cf8043d7f72c · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Curiosity-driven exploration by self-supervised prediction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.820821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.820821Z digest=sha256:230cdfc80c2ec4812becd77b7eb0dfd193185d45c314aca1d63d49d73aaae018

Observation dae2e2b5-728e-483f-a7c1-cdf48f88d5ec · outbound

This paper cites Self-supervised exploration via disagreement.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Self-supervised exploration via disagreement

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.216136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.825285Z digest=sha256:f669fb150b8e64272c1d95bef2fb746a4043e263d20b4ec4caa57b59fe744a62

Observation c253f65d-1884-4552-b4aa-34d7aeb1d127 · outbound

This paper cites Ride: Rewarding impact-driven exploration for procedurally-generated environments.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Ride: Rewarding impact-driven exploration for procedurally-generated environments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.199778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.829753Z digest=sha256:2197550a480c76b417e009cfec1b0577630ca8711405899510fe0585fed2e67c

Observation e7f619b6-92d2-4171-836b-d697f8a232cb · outbound

This paper cites Minihack the planet: A sandbox for open-ended reinforcement learning research.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Minihack the planet: A sandbox for open-ended reinforcement learning research

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.182309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.834609Z digest=sha256:a53eeb96ab98a3fe259d5372906ef861c159e98362e037c3f6d414020bd3115d

Observation 51ccc0d7-1fd0-4e2b-b73b-846d65d96fa9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.839875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.839875Z digest=sha256:54dfbac9b3559d2221953752b391b19554a0a180b3b05bf16a04bb65072fda8c

Observation e019307f-01e9-42f1-9754-8ce787f3b7cb · outbound

This paper cites State entropy maximization with random encoders for efficient exploration.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model State entropy maximization with random encoders for efficient exploration

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.166131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.845382Z digest=sha256:73cbfa898bec40787dbb471ea763c862a8125e7fd1636d26d7484d4d68cb6fb9

Observation 3bd73490-d8c9-4f91-baf4-0b9fe55ed8fc · outbound

This paper cites Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.850563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.850563Z digest=sha256:a55e09c8bc15f2a5b0306325ce8cec09b5ab0f0884837cb636a20b5a40a5b96e

Observation d7c5f284-6475-4019-b3f9-cefc29542115 · outbound

This paper cites Reinforcement learning: An introduction.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning: An introduction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.855932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.855932Z digest=sha256:6a8bc61f43482c5fd97804a9db50633aa2a1b2f272d1a0d05a51a5634923e8cb

Observation 2cef02a3-edf5-43c6-9fc8-4a27ada1fc55 · outbound

This paper cites \# exploration: A study of count-based exploration for deep reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model \# exploration: A study of count-based exploration for deep reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.860690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.860690Z digest=sha256:a38e3a474efe303e8a816c220451683b6ab06f37fcd94321c8b6072b54af25c7

Observation b2bcbb78-6710-494f-8f9a-a84688d5851a · outbound

This paper cites Reinforcement learning with prototypical representations.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Reinforcement learning with prototypical representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.127782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.865517Z digest=sha256:a600d15f169669e371a373d556e47fe9e51aa6d797aca47c6bfb142b28ce720e

Observation ce6a0bb6-ad98-4bdf-8008-b911a70e1984 · outbound

This paper cites Rewarding episodic visitation discrepancy for exploration in reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rewarding episodic visitation discrepancy for exploration in reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.105311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.870168Z digest=sha256:1123efabd159313c5777d2bafdb4494ec46b722192030c46409f297d8f028e02

Observation 3f4e8c4b-f613-45e3-8860-13a13a60dd75 · outbound

This paper cites R \'e nyi state entropy maximization for exploration acceleration in reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model R \'e nyi state entropy maximization for exploration acceleration in reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.089219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.874687Z digest=sha256:5dc0439a723a4a64c30b186227b59e6b367c715de0e348ce42c5bc5d7058389e

Observation f3e1a87a-8f8a-406a-abc1-4777764238e2 · outbound

This paper cites RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.880479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.880479Z digest=sha256:52803d536da366eec5dcc470831973dbcf5adeda0caeef0a23519cd32ca69f5e

Observation 892dd3d4-1673-431a-b114-762c796b0303 · outbound

This paper cites Rllte: Long-term evolution project of reinforcement learning.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Rllte: Long-term evolution project of reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.072848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.886926Z digest=sha256:2a69a8d8dca7990e48ec1bd46e65fd8d307db38ff018059a4ce55f0b22fb98d3

Observation c9f54259-2184-4fac-8134-bcbf9206621d · outbound

This paper cites Noveld: A simple yet effective exploration criterion.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model Noveld: A simple yet effective exploration criterion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:04:32.057202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T17:04:31.891809Z digest=sha256:d8c6be6a77fcec935a15415608bf420c769f8f21b34db49d82b156f97f28b0b0

Observation a0e2ef72-6f7e-4268-8b52-866ea9ee5639 · outbound

This paper cites write newline.

Deep Reinforcement Learning with Hybrid Intrinsic Reward Model write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:04:31.896668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:04:31.896668Z digest=sha256:dcf61e2f85b4d5d3096e7e936adf0820c9111321ccdd6acfdff0d661aaad4994

Pith citing papers

Observation ea666f1f-983f-42cd-9966-94b2d78a1c1d · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.724925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:bbe5f87f478c15d93d78cb04ec5d992e63dfb03be106d61bd2a658a369e45f91