Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:35:56.749428Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2606.18963.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T21:35:56.749428Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ca3b200f-43e2-4324-8d44-ad9a2de8f189 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards OpenAI Gym
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c131fe19-cb3f-42f1-ae28-2d273b00862c · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1edb5864-fa7f-45c6-9663-99f9a9f891c4 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Mastering Diverse Domains through World Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65b74b29-1dce-4bb5-9888-b5668aef2909 · outbound
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2828c0a-d737-4bf8-86c1-abbde782a35b · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Laskin,M.;Yarats,D.;Liu,H.;Lee,K.;Zhan,A.;Lu,K.;Cang,C.; Pinto,L.;andAbbeel,P.2021.URLB:UnsupervisedReinforcement Learning Benchmark
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6872f2-06bc-433a-a3ba-830a52e4cc69 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Proximal Policy Optimization Algorithms
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82580466-26f0-49e7-8125-16e8983d5fc9 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b922f3-844e-479f-b741-7cb72cba76e8 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c95c199c-91b2-4ff6-a7dc-76749bedb1da · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be168cc3-1723-4b87-87da-4155972264a2 · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Puterman, M
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac2da2a3-9878-4488-94dc-71a4b36bb884 · outbound
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a84a3ed1-7807-429a-a2c1-7d46b22aef8c · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fadd1b62-43d5-4422-b18e-e11d56cafc0f · outbound
Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.