Pith. sign in

Paper Citation Record · LEDGER

Position: Deployed Reinforcement Learning should be Continual

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2606.04029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04029 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:52:16.912981Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92ccab8a-f15e-4d28-9476-90046d9d69f8 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Position: Deployed Reinforcement Learning should be Continual Solving Rubik's Cube with a Robot Hand

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:16.602705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:0188d71a3e142cbb094842a972a05ffc93e87b8f284a3e1950b39d42ed8da65c

Observation 1ccad928-0bce-4314-9663-761a3bbb83a8 · outbound

This paper cites Concrete Problems in AI Safety.

Position: Deployed Reinforcement Learning should be Continual Concrete Problems in AI Safety

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:16.611614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:95020ae0f686cc2482b47eb7e76b1e756ca90fc61d0e2ae5228b075b4153cdf8

Observation 7e0f079a-08c9-4156-aba2-3f4718d3a61a · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Position: Deployed Reinforcement Learning should be Continual Dota 2 with Large Scale Deep Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:16.605017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:dd74ba32ac402356058213f6dc1d19cb595603d7f7f2db992b078f0655b5157f

Observation 68d1d2e2-d849-4f29-9a5b-f21463cf11a4 · outbound

This paper cites Step-size Optimization for Continual Learning.

Position: Deployed Reinforcement Learning should be Continual Step-size Optimization for Continual Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:16.617758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:a6487cbf3d7b8eed57506406f867a8e2a2f504bc2887597051325bcb39d23147

Observation b58becd8-7027-4610-86aa-c9d62f85c28e · outbound

This paper cites Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies.

Position: Deployed Reinforcement Learning should be Continual Can Context Bridge the Reality Gap? Sim-to-Real Transfer of Context-Aware Policies

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-28T02:21:36.130719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:fb645d8db4fd036009b04ed93e2e3694da60f0c6eafce17754bcdbda24f383b5

Observation 935d409f-7b00-4f5d-b544-c4f65d3bc493 · outbound

This paper cites Accessed: 2026-01-11.

Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-11

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T15:52:16.912981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:823d5fcaac41a06189abbdf0b6006c4e2ee8c403bd58639e2675a3c84e236c70

Observation 2ae664dc-5123-422c-ba52-fc1bebe5c3c7 · outbound

This paper cites Accessed: 2026-01-28.

Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-28

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T15:52:16.912981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:2bd370193a0654d1dc62057577e518dd0d40f084bff27a01f61019a92366ca49

Observation ec9c1a80-2451-4081-a511-6be0cac50fcf · outbound

This paper cites In-Context Learning can Perform Continual Learning Like Humans.arXiv preprint 2509.22764,.

Position: Deployed Reinforcement Learning should be Continual In-Context Learning can Perform Continual Learning Like Humans.arXiv preprint 2509.22764,

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:16.623050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:a4e7dc85ae5c4eb27d346c9d7e21a7f41ad8be48ade262cfe4a0fbc8c2aae30f

Observation 7963899a-59d5-47d1-8b4b-416028041cfb · outbound

This paper cites Accessed: 2026-01-27.

Position: Deployed Reinforcement Learning should be Continual Accessed: 2026-01-27

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T15:52:16.912981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:eac0b5f73676355ffb286c14f188ceaafc1d03cd7de0201df34e8cf05ff2084f

Observation 7230b9cf-49f5-4521-825d-c050ed5b3d97 · outbound

This paper cites Risk-sensitive Actor-Critic with Static Spectral Risk Measures for Online and Offline Reinforcement Learning.

Position: Deployed Reinforcement Learning should be Continual Risk-sensitive Actor-Critic with Static Spectral Risk Measures for Online and Offline Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:16.628195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:6df941460a15d6b0f05ff953d5027e454ab2ae0d5f4eec8e838b868de2b88313

Observation 5a1215ce-7207-4bd3-b2cd-d9872d7d139d · outbound

This paper cites ualberta.ca/RLAI/rewardhypothesis.

Position: Deployed Reinforcement Learning should be Continual ualberta.ca/RLAI/rewardhypothesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T15:52:16.912981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:98f75890a18bbd866ece464a4fd04b960b51a18447765a28cf7c11def00572ed

Observation 5f31dce2-94e3-45c3-91de-18106c1baa28 · outbound

This paper cites The Alberta Plan for AI Research.

Position: Deployed Reinforcement Learning should be Continual The Alberta Plan for AI Research

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:56:16.625429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:dc49eda9d27b0efb847c9a7d2d5275c96b11adc4fe0ca791c7ad5daf61ed9e4a

Observation ad5cb1a3-aa96-40c4-bc86-1be31eb72f77 · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Position: Deployed Reinforcement Learning should be Continual Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:56:16.614234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:8424111ffc8b5241e8b913f57b0b25ae6177184f20724de7c6b0e47447f4357e

Observation 57da014b-a1ab-4e2f-860b-dd73df5d1e51 · outbound

This paper cites On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs.

Position: Deployed Reinforcement Learning should be Continual On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:16.608428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:2fa30735210adbb65ace36420110e4172a2eb937ff2a1e2cd9b79ddd0760f634

Observation 62716def-181e-40de-9008-4a5dd09ed3e7 · outbound

This paper cites The damping coefficient bt grows in a noisily quadratic manner over time, as shown in Figure 3 (Gaussian noise σ= 0.02 ).

Position: Deployed Reinforcement Learning should be Continual The damping coefficient bt grows in a noisily quadratic manner over time, as shown in Figure 3 (Gaussian noise σ= 0.02 )

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T15:52:16.912981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:52:16.912981Z digest=sha256:96bb970aedeac970ef0b77e87acb5a9a8f660cc823dcaa85a2cfca557aa86b37

Pith citing papers

No inbound Pith citation observations are available.