Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T10:49:33.177939Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:1908.10479.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T10:49:33.177939Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:03:42.755190Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T00:03:43.422793Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 48ef1794-c2dd-4d7b-8323-0adabe0aecef · outbound
Exploration-Enhanced POLITEX POLITEX : Regret bounds for policy iteration using expert prediction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fd36366b-2485-42d4-8228-1936d8644c7c · outbound
Exploration-Enhanced POLITEX Model-free linear quadratic control via reduction to expert prediction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 01862f8c-70b0-4346-87cb-27b89e6e07ae · outbound
Exploration-Enhanced POLITEX Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2818fad7-6aa2-45d7-a873-dd6db7f25f31 · outbound
Exploration-Enhanced POLITEX Multi-step Reinforcement Learning: A Unifying Algorithm
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa4bafc-96cb-423e-af40-3a321bfe4cf3 · outbound
Exploration-Enhanced POLITEX Minimax regret bounds for reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e70bac70-f61a-4cce-a353-bf5fc3d8656f · outbound
Exploration-Enhanced POLITEX Approximate policy iteration: A survey and some new methods
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4dd4e0e3-c6b0-431f-8048-9e3e111e0541 · outbound
Exploration-Enhanced POLITEX Temporal differences-based policy iteration and applications in neuro-dynamic programming
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4dcfc5f9-2f35-41fd-9c8b-de263ca7e93f · outbound
Exploration-Enhanced POLITEX Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92f9c96b-aeee-4646-b179-6d00e11ed4bb · outbound
Exploration-Enhanced POLITEX Regularized policy iteration with nonparametric function spaces
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 150164ba-4d00-40dc-89bf-5c3c9b507d97 · outbound
Exploration-Enhanced POLITEX Off-policy learning with eligibility traces: A survey
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3911b1cf-d4f8-4e78-8fd4-a1bfef26719b · outbound
Exploration-Enhanced POLITEX Provably efficient maximum entropy exploration
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6f5641a9-04f6-4347-af17-f545a49de972 · outbound
Exploration-Enhanced POLITEX Distributed Prioritized Experience Replay
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119329cf-cced-4c11-b5a4-9a92a66ce64e · outbound
Exploration-Enhanced POLITEX Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1ef180e9-b097-4d0b-a139-cc258494d1c7 · outbound
Exploration-Enhanced POLITEX Finite-sample analysis of least-squares policy iteration
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 32fe44af-31a9-4e3d-94b5-822ac92be12d · outbound
Exploration-Enhanced POLITEX Regularized off-policy TD -learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf4622c3-436d-47f7-851a-5806330e576a · outbound
Exploration-Enhanced POLITEX Finite-sample analysis of proximal gradient TD algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d696bc50-07a6-454b-861a-9b17281ec7d5 · outbound
Exploration-Enhanced POLITEX Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f55adedd-70ee-49cc-8135-33279b812910 · outbound
Exploration-Enhanced POLITEX Human-level control through deep reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1c481485-fdc4-41c5-9296-3fd39b20ef00 · outbound
Exploration-Enhanced POLITEX Scale-free algorithms for online linear optimization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b301f1c8-0099-461c-8985-b71b77f20c04 · outbound
Exploration-Enhanced POLITEX Generalization and exploration via randomized value functions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 217eb29d-4adb-4868-bf0c-6afc9121e184 · outbound
Exploration-Enhanced POLITEX Puterman
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c59fedf-5489-4d13-a167-40afcd39d61b · outbound
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d37fd51-0cf9-4b9c-8ed4-175cec303a1d · outbound
Exploration-Enhanced POLITEX Adaptive confidence and adaptive curiosity
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a7a1475-8128-41e0-99f1-85ec335a5981 · outbound
Exploration-Enhanced POLITEX Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6c5b7489-d6d7-4573-b0c7-5becb7ddfad4 · outbound
Exploration-Enhanced POLITEX Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 358f5f31-fe06-454e-91fa-60d182157560 · outbound
Exploration-Enhanced POLITEX Learning to predict by the methods of temporal differences
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d999fd21-c414-4317-8ded-21298ed57e68 · outbound
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49b2a35-4d3b-4464-89e3-4c7f54b60dfb · outbound
Exploration-Enhanced POLITEX Active exploration in dynamic environments
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db1c6b24-923a-423c-b1ae-af0bfae17c19 · outbound
Exploration-Enhanced POLITEX Tsitsiklis and Benjamin Van Roy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2a7ef33-16d8-429d-8f71-556d9cec8dea · outbound
Exploration-Enhanced POLITEX Tsitsiklis and Benjamin Van Roy
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 70a53df3-c5cd-4e5c-b051-a6ab4c4ea169 · outbound
Exploration-Enhanced POLITEX Deep Reinforcement Learning with Double Q-learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fd6f24-10ca-4fca-93a6-421e704b76f8 · outbound
Exploration-Enhanced POLITEX Dueling Network Architectures for Deep Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da53bf8a-c7da-4036-8103-49e8dd40661e · outbound
Exploration-Enhanced POLITEX Convergence of least squares temporal difference methods under general conditions
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 876a06b6-9a15-47ca-86ec-f3e88429d4ed · outbound
Exploration-Enhanced POLITEX Convergence results for some temporal difference methods based on least squares
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0e71d808-5563-48ff-9b46-a24dd1ae57c4 · outbound
Exploration-Enhanced POLITEX Error bounds for approximations from projected linear equations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8773c8a5-d5a1-4949-925b-58fe74f2e113 · inbound
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation Exploration-Enhanced POLITEX
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.