Pith. sign in

Paper Citation Record · LEDGER

Exploration-Enhanced POLITEX

As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:1908.10479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10479 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:49:33.177939Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:03:42.755190Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:03:43.422793Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48ef1794-c2dd-4d7b-8323-0adabe0aecef · outbound

This paper cites POLITEX : Regret bounds for policy iteration using expert prediction.

Exploration-Enhanced POLITEX POLITEX : Regret bounds for policy iteration using expert prediction

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.586541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.039204Z digest=sha256:6b672b83eb3f51317a90488e2405d23d7e723707dd4baa66483bff04bf4be233

Observation fd36366b-2485-42d4-8228-1936d8644c7c · outbound

This paper cites Model-free linear quadratic control via reduction to expert prediction.

Exploration-Enhanced POLITEX Model-free linear quadratic control via reduction to expert prediction

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.575488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.043436Z digest=sha256:fd1a20afb091e9963f6de8e81ada1751acf418a19ce9612e54a37fd518f5b44f

Observation 01862f8c-70b0-4346-87cb-27b89e6e07ae · outbound

This paper cites Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path.

Exploration-Enhanced POLITEX Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.564109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.047195Z digest=sha256:5881158b2b70ecee597458adfdaec47df69240cf8d25e9708413a12f0948f480

Observation 2818fad7-6aa2-45d7-a873-dd6db7f25f31 · outbound

This paper cites Multi-step Reinforcement Learning: A Unifying Algorithm.

Exploration-Enhanced POLITEX Multi-step Reinforcement Learning: A Unifying Algorithm

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.051233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.051233Z digest=sha256:177d9cb5bb753450a8f8e61c911423e31b5454d136332691acb0c2cae949958a

Observation 7fa4bafc-96cb-423e-af40-3a321bfe4cf3 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Exploration-Enhanced POLITEX Minimax regret bounds for reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.553879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.055695Z digest=sha256:ac5a655e6751268c38c126dfbc532b3006b5129a7b3b41092b9b2e1304f5d927

Observation e70bac70-f61a-4cce-a353-bf5fc3d8656f · outbound

This paper cites Approximate policy iteration: A survey and some new methods.

Exploration-Enhanced POLITEX Approximate policy iteration: A survey and some new methods

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.543068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.059970Z digest=sha256:daef30411d39ed2887e3b296de0561c1859f1bf56302530dcbcd2211600677ef

Observation 4dd4e0e3-c6b0-431f-8048-9e3e111e0541 · outbound

This paper cites Temporal differences-based policy iteration and applications in neuro-dynamic programming.

Exploration-Enhanced POLITEX Temporal differences-based policy iteration and applications in neuro-dynamic programming

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.532219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.064174Z digest=sha256:41b1e66b9c3ebc4a39c35782610609d946bff41d6ca3a2216531301a90a723b8

Observation 4dcfc5f9-2f35-41fd-9c8b-de263ca7e93f · outbound

This paper cites an unresolved cited work.

Exploration-Enhanced POLITEX Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:49:33.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.068513Z digest=sha256:e6b009f6640299642c219469c18ec28a19c015bc8781b6b17b20b859a0978576

Observation 92f9c96b-aeee-4646-b179-6d00e11ed4bb · outbound

This paper cites Regularized policy iteration with nonparametric function spaces.

Exploration-Enhanced POLITEX Regularized policy iteration with nonparametric function spaces

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.509328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.072931Z digest=sha256:cd76a271959dde1a649c8e6559c5deeae2ff5f7e6a214afa7a5e5db6ce69530e

Observation 150164ba-4d00-40dc-89bf-5c3c9b507d97 · outbound

This paper cites Off-policy learning with eligibility traces: A survey.

Exploration-Enhanced POLITEX Off-policy learning with eligibility traces: A survey

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.497765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.077195Z digest=sha256:55f95d1cdf4eb17c1b8a34d3877fed0baf59b80a37ae77e036349e048578251e

Observation 3911b1cf-d4f8-4e78-8fd4-a1bfef26719b · outbound

This paper cites Provably efficient maximum entropy exploration.

Exploration-Enhanced POLITEX Provably efficient maximum entropy exploration

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.486557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.081374Z digest=sha256:527d3b7c5d969d521fb51786610daaadfed9823d4aa1686a871b9328592b4bd5

Observation 6f5641a9-04f6-4347-af17-f545a49de972 · outbound

This paper cites Distributed Prioritized Experience Replay.

Exploration-Enhanced POLITEX Distributed Prioritized Experience Replay

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.086022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.086022Z digest=sha256:4bd9961529042b1aaec3d59c2d5f414c441473922013008ad3e2a1fa53b44798

Observation 119329cf-cced-4c11-b5a4-9a92a66ce64e · outbound

This paper cites an unresolved cited work.

Exploration-Enhanced POLITEX Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:49:33.474588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.089703Z digest=sha256:d6f10d6af9151108d2307eecaa5a37d3a457981ff962f6eb9e0c4086419447bd

Observation 1ef180e9-b097-4d0b-a139-cc258494d1c7 · outbound

This paper cites Finite-sample analysis of least-squares policy iteration.

Exploration-Enhanced POLITEX Finite-sample analysis of least-squares policy iteration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.463846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.093664Z digest=sha256:28534f38a714244df8c1caa0a9de7916546073cc1e50bdac32b50881d90bbeb7

Observation 32fe44af-31a9-4e3d-94b5-822ac92be12d · outbound

This paper cites Regularized off-policy TD -learning.

Exploration-Enhanced POLITEX Regularized off-policy TD -learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.451767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.097990Z digest=sha256:aa3cffde939b6185a7430a5e7abc14e38c674110db0d6ac2515e7f6e9f9ae789

Observation bf4622c3-436d-47f7-851a-5806330e576a · outbound

This paper cites Finite-sample analysis of proximal gradient TD algorithms.

Exploration-Enhanced POLITEX Finite-sample analysis of proximal gradient TD algorithms

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.439104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.101809Z digest=sha256:e51ca95802d5a4667e5e5ea84c02a06e838188ba82dd859595fac16e297bbd43

Observation d696bc50-07a6-454b-861a-9b17281ec7d5 · outbound

This paper cites an unresolved cited work.

Exploration-Enhanced POLITEX Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:49:33.427193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.105720Z digest=sha256:bbbad957b73f700e7a5de8a2c4f63fb6e115f2c672f70e7dee8d6f4be66d7e5c

Observation f55adedd-70ee-49cc-8135-33279b812910 · outbound

This paper cites Human-level control through deep reinforcement learning.

Exploration-Enhanced POLITEX Human-level control through deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.416139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.109950Z digest=sha256:f2722ba9b40b41f6fcaf5794f0b8307723d3f91d9081a44eb6bb588c842071f1

Observation 1c481485-fdc4-41c5-9296-3fd39b20ef00 · outbound

This paper cites Scale-free algorithms for online linear optimization.

Exploration-Enhanced POLITEX Scale-free algorithms for online linear optimization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.405308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.113990Z digest=sha256:55d9e32977aaee55cfaabb04636eb14fa9784d22cb136058fc3ad7a9567d8fa7

Observation b301f1c8-0099-461c-8985-b71b77f20c04 · outbound

This paper cites Generalization and exploration via randomized value functions.

Exploration-Enhanced POLITEX Generalization and exploration via randomized value functions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.393205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.118323Z digest=sha256:851d2b0af9e55fb729f9b23fd52defe99ac37d6c96eed7cd4565f6a6531ba65c

Observation 217eb29d-4adb-4868-bf0c-6afc9121e184 · outbound

This paper cites Puterman.

Exploration-Enhanced POLITEX Puterman

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.381734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.122121Z digest=sha256:a135074db6f97970bf6207c50bf377b92d9decdfaf7fdc164c650ae2c97a2e93

Observation 5c59fedf-5489-4d13-a167-40afcd39d61b · outbound

This paper cites Prioritized Experience Replay.

Exploration-Enhanced POLITEX Prioritized Experience Replay

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.126218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.126218Z digest=sha256:c0bd7c258e7d1a57662f924b4d68dd744ce1c56cf88e414e8982ecc20fe20dfe

Observation 7d37fd51-0cf9-4b9c-8ed4-175cec303a1d · outbound

This paper cites Adaptive confidence and adaptive curiosity.

Exploration-Enhanced POLITEX Adaptive confidence and adaptive curiosity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.370843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.130160Z digest=sha256:15cf15f861ae5f3002bdf0719c926a55478f17610cf74141296ced83976a7d40

Observation 6a7a1475-8128-41e0-99f1-85ec335a5981 · outbound

This paper cites Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L.

Exploration-Enhanced POLITEX Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.359154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.134516Z digest=sha256:e6f0464c35edd0a3fb0aa01231c4bca0ba9c6f9cd6a4e39c1283116386c683fa

Observation 6c5b7489-d6d7-4573-b0c7-5becb7ddfad4 · outbound

This paper cites an unresolved cited work.

Exploration-Enhanced POLITEX Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:49:33.348019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.138301Z digest=sha256:d9f6e64247120c512c7087b3849916d1680b155bc58afd2e16d9576f3c72b76f

Observation 358f5f31-fe06-454e-91fa-60d182157560 · outbound

This paper cites Learning to predict by the methods of temporal differences.

Exploration-Enhanced POLITEX Learning to predict by the methods of temporal differences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.142285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.142285Z digest=sha256:2a906c5bbdd41fe84b4b12d0d8d2bbda4c9512286db2134fe458e7b42c5a4e19

Observation d999fd21-c414-4317-8ded-21298ed57e68 · outbound

This paper cites DeepMind Control Suite.

Exploration-Enhanced POLITEX DeepMind Control Suite

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.146230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.146230Z digest=sha256:c780bd23ac6def1def05f25795c5b6447db3e01776462a14d77ad6e5733ac686

Observation d49b2a35-4d3b-4464-89e3-4c7f54b60dfb · outbound

This paper cites Active exploration in dynamic environments.

Exploration-Enhanced POLITEX Active exploration in dynamic environments

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.328788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.150267Z digest=sha256:92ceb8950a3eb2b0850a29a380eac8cb6de3a6c3ed248ec981bd1b7d8456e225

Observation db1c6b24-923a-423c-b1ae-af0bfae17c19 · outbound

This paper cites Tsitsiklis and Benjamin Van Roy.

Exploration-Enhanced POLITEX Tsitsiklis and Benjamin Van Roy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.315837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.154018Z digest=sha256:9a86c16689f84f3e27528f4eb5073cb43f9eddc402d107c605b903ed7c267f6e

Observation f2a7ef33-16d8-429d-8f71-556d9cec8dea · outbound

This paper cites Tsitsiklis and Benjamin Van Roy.

Exploration-Enhanced POLITEX Tsitsiklis and Benjamin Van Roy

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.304244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.157670Z digest=sha256:9854e6f1c014b37e2c1613e73a9028d3d2fdae045d3cf25932d1fea576becc50

Observation 70a53df3-c5cd-4e5c-b051-a6ab4c4ea169 · outbound

This paper cites Deep Reinforcement Learning with Double Q-learning.

Exploration-Enhanced POLITEX Deep Reinforcement Learning with Double Q-learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.161700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.161700Z digest=sha256:46cbbddc667fcc07fd1c1ccdf14ff35d7f5cd5180cbc2f4b5260effc11bab11a

Observation 63fd6f24-10ca-4fca-93a6-421e704b76f8 · outbound

This paper cites Dueling Network Architectures for Deep Reinforcement Learning.

Exploration-Enhanced POLITEX Dueling Network Architectures for Deep Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T10:49:33.165609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:49:33.165609Z digest=sha256:f696c77118f8d43326b6e0c013c532e14b679157015acb66b9f8cc9c1bee4b6a

Observation da53bf8a-c7da-4036-8103-49e8dd40661e · outbound

This paper cites Convergence of least squares temporal difference methods under general conditions.

Exploration-Enhanced POLITEX Convergence of least squares temporal difference methods under general conditions

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.293128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.169650Z digest=sha256:41ba622f7bcf9523a4f24232c8b48900c5f474eac0faef1a2f5f1d7766aa1d14

Observation 876a06b6-9a15-47ca-86ec-f3e88429d4ed · outbound

This paper cites Convergence results for some temporal difference methods based on least squares.

Exploration-Enhanced POLITEX Convergence results for some temporal difference methods based on least squares

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.281903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.173306Z digest=sha256:5e228429515157a3d6ae05e74404050efdd1847767d95562c77ae6e20ede878e

Observation 0e71d808-5563-48ff-9b46-a24dd1ae57c4 · outbound

This paper cites Error bounds for approximations from projected linear equations.

Exploration-Enhanced POLITEX Error bounds for approximations from projected linear equations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:49:33.270198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T10:49:33.177939Z digest=sha256:dc4afc6daf74a284c1f30f8a5b9c8755c1cd4a09d51c208795a1923dcbc491fa

Pith citing papers

Observation 8773c8a5-d5a1-4949-925b-58fe74f2e113 · inbound

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation cites this paper.

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation Exploration-Enhanced POLITEX

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:03:43.433698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T00:03:42.755190Z digest=sha256:a9ccdec1c254085d1112a8d868a29127a913962b37257825b77777cb717dcdf7