Pith. sign in

Paper Citation Record · LEDGER

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments

As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2507.00030.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00030 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:12:49.665075Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1a3798-0d25-4f56-b343-8670aff436f4 · outbound

This paper cites Bellemare and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Bellemare and others

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.530869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.475803Z digest=sha256:f1e94aa057933c3984c1a7b084add3f4e9ea36d6b8e37325d965585be31daaae

Observation 3567d400-933f-469b-b9bb-2ca1b7f304a1 · outbound

This paper cites Frame skip is a powerful parameter for learning to play Atari.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Frame skip is a powerful parameter for learning to play Atari

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.398444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.543454Z digest=sha256:6f6ab7249b459907b70783b752fc9b9e0b135143f78ce7ce3b3539dab0237718

Observation 335db7ef-2e9f-4867-8c0d-da5f348a8d89 · outbound

This paper cites Gilbert and Timothy D.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Gilbert and Timothy D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.272890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.647439Z digest=sha256:92bf6f67b5775739171b92a0bfc81be73900f13352a001941287b0c46783de7e

Observation a555fbe8-65b9-41fd-8310-f78551845d19 · outbound

This paper cites Lakshminarayanan and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Lakshminarayanan and others

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:51.099339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.777519Z digest=sha256:cf9a8e16b2ea3271b56ca3c569e95689a7eaea31ce92b0a5810b0b1e42e32294

Observation 491d555f-1113-4195-b209-78488a0a87e3 · outbound

This paper cites A contextual-bandit approach to personalized news article recommendation.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A contextual-bandit approach to personalized news article recommendation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.919715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.881659Z digest=sha256:61e60a5d7a9940363ad702445fcab275d3687045eaef7862c9455d233f686669

Observation 6bf59c52-fba4-4fd7-a67d-3803fe9168ce · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Human-level control through deep reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.740473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:48.954347Z digest=sha256:e16ff495de91a99108fcc13b5ddbdbf482c66d8f941863d5c580961b6db34749

Observation b9ff5d85-7e5b-45f0-82f3-613ccedff340 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Asynchronous methods for deep reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.590134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.014473Z digest=sha256:52137577bebed4d6f690b3fe070b60da3af9250822ea2d16a844b14ce6f102b2

Observation 0e8db6ed-34d1-402c-9815-b9aabcb59659 · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Mastering the game of Go with deep neural networks and tree search

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.419941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.154324Z digest=sha256:c3c801ecc8405cbc5e96da9f97a4e82e38462e38a96071a438b7b1abf87527c9

Observation 26e464ee-730c-4a27-95b4-e2089cef3724 · outbound

This paper cites Sutton and others.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Sutton and others

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:12:50.250126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.316626Z digest=sha256:4add38fa7e5e94368ed4f0628a7124c0ee3f6abca95ab9bd3fb9ffe46e00fc44

Observation 4aabe205-466b-4c2b-a5c5-00452fb8a38b · outbound

This paper cites Effect of scalar leptoquarks on the rare decays of B_s meson.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments Effect of scalar leptoquarks on the rare decays of B_s meson

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:49.836013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.524676Z digest=sha256:2dce0cf9f3788bba098232efbdcc2ffc2224190e18fef6e05f84e697e55927f5

Observation 9daf9023-c8d9-4a86-93e8-5808381c8ac9 · outbound

This paper cites A gradient estimate for nonlocal minimal graphs.

Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments A gradient estimate for nonlocal minimal graphs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:12:50.085655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-07T00:12:49.665075Z digest=sha256:0a7c9413fcb081260452d5dff6c561c1ec99a05b8ea9c5e6e4a107da8089b036

Pith citing papers

No inbound Pith citation observations are available.