Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:20:01.500799Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2501.05501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:20:01.500799Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b0f7a53-54a8-4bcd-b3be-61cd51e7b03c · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Never Give Up: Learning Directed Exploration Strategies
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3bc92a-e2ae-4d1f-ae4b-7f916ae0e398 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents OpenAI Gym
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f02a0f-43f7-4ea9-9b63-4ec2c605b1c2 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Brown and T
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5ad7fb-b1d6-48a7-8f75-2530a001e4cb · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Brown and T
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4387aa5f-df0a-4520-a506-5d6d572db801 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Dietterich
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 353d1950-63f9-4b0b-bacd-0b98ae2168e8 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Deep Recurrent Q-Learning for Partially Observable MDPs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7134c656-60c3-436d-9273-2643819e0329 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Hochreiter and J
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22847d35-5279-4f1a-92bd-16e6a63b1977 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09188085-a1a5-41f1-83ea-699c165066e5 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Jaakkola, M
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af2a37f1-e073-4d29-9beb-e8beaba00405 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Juozapaitis, A
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59d64928-dbf3-4040-9f88-f8c5abfd838f · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Karlsson
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed44d714-c9fe-4cc4-a7b9-e42c80817f31 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Playing Atari with Deep Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca63208-5a39-437a-8245-0b4eef556b3f · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8416c1d6-f018-4c69-8693-b0bf36f10b97 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb072f77-52b8-4b67-bafb-2735f4b3cea1 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd345702-7763-42ca-b332-6ecaaa48dbed · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f324169c-03ed-4096-8f9f-fd7393fb20ce · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 55632a50-b590-42cd-b58f-ec5c5c28b98a · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Hybrid Reward Architecture for Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84584a4-f9ca-4dae-b5f8-b56e23fd1925 · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Vinyals, I
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d720c3d6-e188-43ad-8953-50f1a06cea0e · outbound
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.