Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:12:49.665075Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.00030.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:12:49.665075Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
11 of 11 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e1a3798-0d25-4f56-b343-8670aff436f4 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Bellemare and others
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3567d400-933f-469b-b9bb-2ca1b7f304a1 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Frame skip is a powerful parameter for learning to play Atari
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 335db7ef-2e9f-4867-8c0d-da5f348a8d89 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Gilbert and Timothy D
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a555fbe8-65b9-41fd-8310-f78551845d19 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Lakshminarayanan and others
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 491d555f-1113-4195-b209-78488a0a87e3 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A contextual-bandit approach to personalized news article recommendation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6bf59c52-fba4-4fd7-a67d-3803fe9168ce · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Human-level control through deep reinforcement learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b9ff5d85-7e5b-45f0-82f3-613ccedff340 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Asynchronous methods for deep reinforcement learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e8db6ed-34d1-402c-9815-b9aabcb59659 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Mastering the game of Go with deep neural networks and tree search
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 26e464ee-730c-4a27-95b4-e2089cef3724 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Sutton and others
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4aabe205-466b-4c2b-a5c5-00452fb8a38b · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Effect of scalar leptoquarks on the rare decays of B_s meson
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9daf9023-c8d9-4a86-93e8-5808381c8ac9 · outbound
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A gradient estimate for nonlocal minimal graphs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.