Pith. sign in

Paper Citation Record · LEDGER

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2606.18963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.18963 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T21:35:56.749428Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca3b200f-43e2-4324-8d44-ad9a2de8f189 · outbound

This paper cites OpenAI Gym.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards OpenAI Gym

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.821564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:ca14f4d9cc8c303244fce559e3cd80436b66ec61c5085e0b28407131f101235b

Observation c131fe19-cb3f-42f1-ae28-2d273b00862c · outbound

This paper cites an unresolved cited work.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:165f062c2d0243651c31e20087f5487814408e1d2715e364a556a60ac1a57a03

Observation 1edb5864-fa7f-45c6-9663-99f9a9f891c4 · outbound

This paper cites Mastering Diverse Domains through World Models.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Mastering Diverse Domains through World Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:59:06.813938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:4317e557f93d5a51ba48a31307a81ca1db1e5a62b262464cc993ec7613bd6138

Observation 65b74b29-1dce-4bb5-9888-b5668aef2909 · outbound

This paper cites InNeurIPS.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards InNeurIPS

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:cef27ffcca730d3d57f5b3ead633f5e4f00e2ef92d4ef54cb61ad2c65d4428ff

Observation d2828c0a-d737-4bf8-86c1-abbde782a35b · outbound

This paper cites Laskin,M.;Yarats,D.;Liu,H.;Lee,K.;Zhan,A.;Lu,K.;Cang,C.; Pinto,L.;andAbbeel,P.2021.URLB:UnsupervisedReinforcement Learning Benchmark.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Laskin,M.;Yarats,D.;Liu,H.;Lee,K.;Zhan,A.;Lu,K.;Cang,C.; Pinto,L.;andAbbeel,P.2021.URLB:UnsupervisedReinforcement Learning Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:06dbdb3fc4f9354bd55babb1ae2944b037b128ad00290d5815afb947baaed5e0

Observation da6872f2-06bc-433a-a3ba-830a52e4cc69 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Proximal Policy Optimization Algorithms

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.825532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:276844ab598866d0bfd00096c496f46f4b73d6339cc89ef5168f28e20002e6e0

Observation 82580466-26f0-49e7-8125-16e8983d5fc9 · outbound

This paper cites an unresolved cited work.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:4d40eb4adea60b9ecf4cf2b1b1101e33f96f92a6d36659b13109fd94f9b6c42b

Observation c2b922f3-844e-479f-b741-7cb72cba76e8 · outbound

This paper cites Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.817201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:d6abec7b6517c32b14248b976f6ae892f140f726dabca1ddbb1fe15bcd1547bf

Observation c95c199c-91b2-4ff6-a7dc-76749bedb1da · outbound

This paper cites Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:06.801171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:10cc84fcd1426017488c36f2ce35dc3b3e9b7e52be9f67bfc43ee6ab573986df

Observation be168cc3-1723-4b87-87da-4155972264a2 · outbound

This paper cites Puterman, M.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Puterman, M

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:2d4067b54d6ec6681e351a3623c665d0df36e737c92013c178cba3e3d54ea242

Observation ac2da2a3-9878-4488-94dc-71a4b36bb884 · outbound

This paper cites Sutton, R.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Sutton, R

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T21:35:56.749428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:00ce20a34fc4d349f8ddad66fb307f0d53ab935ca82ed086916acf0eed8b0f0d

Observation a84a3ed1-7807-429a-a2c1-7d46b22aef8c · outbound

This paper cites Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Intrinsic Rewards for Exploration without Harm from Observational Noise: A Simulation Study Based on the Free Energy Principle

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.808701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:d7bf7d44b65c06d0a76d3831662f1eb7324b0e96bcdbe59e2a20b51cda582d09

Observation fadd1b62-43d5-4422-b18e-e11d56cafc0f · outbound

This paper cites RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.805552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:3e62fc547b05b7606a6a4e6f86c824cf1c4ba5cd9e33b88e2429f69e21e156a6

Pith citing papers

No inbound Pith citation observations are available.