Pith. sign in

Paper Citation Record · LEDGER

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.17201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17201 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:54:01.147324Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7151ce36-21b7-4331-9ea6-7279ad572995 · outbound

This paper cites Al Marjani and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Al Marjani and A

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.493874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.493874Z digest=sha256:f2bb9ebdfb6ded5ec95c71ae429d1a97887b2b540a8cb9028735e6834faa6d60

Observation 06e02925-65d5-44b7-a941-788000b5515e · outbound

This paper cites Al Marjani, A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Al Marjani, A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.529153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.529153Z digest=sha256:960a8cf3122bf2a7bccf92b953a667773cbb96ae319eaff0d2f5af9171e4a7ae

Observation eb8f7d5f-4cd3-4159-bab6-483cb77255c7 · outbound

This paper cites Policy Testing in Markov Decision Processes.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Policy Testing in Markov Decision Processes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.581627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.581627Z digest=sha256:d68b028d596f103ec402094d407f130d14919952d8db8a4a4de6116605fc80b3

Observation 5db72b6e-b558-44fb-84bb-66656e05f635 · outbound

This paper cites Berge.Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces, and Convexity.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Berge.Topological Spaces: Including a Treatment of Multi-Valued Functions, Vector Spaces, and Convexity

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.645968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.645968Z digest=sha256:292c513849a6c8d5793a39b87016bd2b4fb7badabce1477e469e45b91b51b499

Observation d909eea6-b31c-4031-a7f4-2b717132859b · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning The regret lower bound for communicating Markov Decision Processes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.706383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.706383Z digest=sha256:0354085614833c553e580f80fa986e86f22867afac1ecd653e85b88d85a69efb

Observation 57780a37-c1e1-4079-bbe6-c2a72904419b · outbound

This paper cites Boucheron, G.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Boucheron, G

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.770229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.770229Z digest=sha256:c5e83c842f73aa5a58711be48bf54be0d741e4ef78ddc6a74ee2ccd7920aaab8

Observation 35b958e0-4312-4ee4-8b38-d8bd4a94e922 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.826541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.826541Z digest=sha256:8e5649bce4d09bcf28c2e8c6dd07269a2ef93a4ebdf44dda79ba736751eaf82c

Observation cc0fdbbe-2b6f-468b-9fab-af3a2ff57e66 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-01T18:59:12.738112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-01T18:53:58.871081Z digest=sha256:983582205b22a1da098ab024baa1e949b5d8c9d3b4c8ab1c37d5d9e609961474

Observation c7876929-fb53-45a4-9151-039a347cdc81 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:58.945774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:58.945774Z digest=sha256:ceb825d128f8b726feeede813ca1fa0a287378551159ef4b905b88d4b0076cc7

Observation 89863d45-3678-4211-85ce-8f2e6f2a8cd6 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.009892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.009892Z digest=sha256:bc4e72a001ba85188b61f8e30d8f362cdda762d800da0e13e56f32486f238018

Observation 8f0bcc96-f663-41c7-b677-5d9384688533 · outbound

This paper cites Degenne and W.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Degenne and W

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.116553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.116553Z digest=sha256:85c128db90b13382e450116db69bf6ed08a670f6265897cd9c31e3cdce661e6e

Observation b9d2103a-99e4-478c-a38f-eebcc5d336bb · outbound

This paper cites Degenne, W.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Degenne, W

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.167718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.167718Z digest=sha256:77102b84a8ae02d8c4e0abfc14648bd9c5610d82f12432f28668001dff0a9dd6

Observation 3d740a2c-859d-47ca-859f-242de7a9d63c · outbound

This paper cites Garivier and E.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Garivier and E

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.262719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.262719Z digest=sha256:461c3560f5be3e3ba7f61ef0bb0082b7ed0143fc85cff7852ad899b15cee0980

Observation 03a5005e-b5c8-4b51-affd-a8ca7f56596f · outbound

This paper cites Thresholding Bandit for Dose-ranging: The Impact of Monotonicity.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Thresholding Bandit for Dose-ranging: The Impact of Monotonicity

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.352381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.352381Z digest=sha256:7418e3cc57b9a8355bafdab6ebcbd9ff0ee02fd1a6a5bb36c86958152d84aa86

Observation ea5a2da7-1be2-42c6-90c7-e54f0ed970c6 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.442412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.442412Z digest=sha256:d77fd4d030f2925b33a3e14be6ff9b77a1cea0ffe69dc2eb0bbe81a82763923f

Observation 0932b424-0528-462e-8a90-dcd389e59d9f · outbound

This paper cites Jonsson, E.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jonsson, E

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.452766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.452766Z digest=sha256:e167f4d5870b3c5115d3a1fe369ba1093d60582affdc1f006244bad3f3f0fb15

Observation 1d2e93fe-fe6a-480f-ab5a-d8737b0ea7cb · outbound

This paper cites Jourdan and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jourdan and A

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.511504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.511504Z digest=sha256:354b376457307581f36c1e6e5403a4802004d223fa0356fe852f5cfb51c047ae

Observation 1f0be86c-19be-4cb5-8176-6cacd2e5c25f · outbound

This paper cites Jourdan, R.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Jourdan, R

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.578201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.578201Z digest=sha256:2e07c43a6c37ccfdb72b7e9b108269791f18dc168051ed056476222193870337

Observation 57eccf5b-b2b7-4de4-970c-f440032526d9 · outbound

This paper cites Kaufmann, P.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Kaufmann, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.692268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.692268Z digest=sha256:d05dfcc6e979b33b4e828e533748e19bada1035d0f2bbcde224bb5febfc10193

Observation d906c45c-3b27-448f-be8d-5343fa621c18 · outbound

This paper cites Lazzaro and C.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Lazzaro and C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.790934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.790934Z digest=sha256:41767316880affff01b298c8bc322e789a7eeea0ba5844136e065058b944d16a

Observation 842bcfb4-453a-442e-882f-d3e4cf19f9a7 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.854229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.854229Z digest=sha256:441c765d7c28cd59a6927cda6aabd12050b4cfb9bd0e27f503ddfb6801e53e84

Observation dc50e33e-8360-45a6-9365-1d9cb6cbf796 · outbound

This paper cites Poiani, M.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Poiani, M

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:53:59.941681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:53:59.941681Z digest=sha256:089c569eeb5757c42647670038de92b7003c88b78d4d12971805c6773833148f

Observation 10d17296-4c7f-4280-b064-9b74ca0f8251 · outbound

This paper cites Poiani, M.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Poiani, M

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.031409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.031409Z digest=sha256:2459e6205013045bb0f855533669962e09cf8799a96b707b3f5804e565ea2110

Observation a3557330-acb1-4377-9e61-ef785a6d2b92 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.132912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.132912Z digest=sha256:c48d45362cb65b5a3cdabde0474a5168f81b4a83118968db08765387accaa59f

Observation 661319d6-b19e-4270-8ce8-866f141e0195 · outbound

This paper cites Adaptive Exploration for Multi-Reward Multi-Policy Evaluation.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.193317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.193317Z digest=sha256:8f681d88f8796e7216476613c7f86ea88c263f38b7a6807d5aca9d6b4788dbad

Observation 1b4326d2-0013-4adf-a545-e70733671d2e · outbound

This paper cites Russo and A.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Russo and A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.274126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.274126Z digest=sha256:c36ccd0ce876db9d68252b6828be7ec44e6cc871a83ac1eba1a581b1b76fe5f7

Observation 224278e4-ab14-4781-9513-404ffd466581 · outbound

This paper cites Russo and F.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Russo and F

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.432674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.432674Z digest=sha256:1ffb4273125f2308e119bb4707be6566856fe3822cc62eb6f02af52b4cb3eb9c

Observation f0043074-c72f-4304-8351-3d4601eaf280 · outbound

This paper cites Pure Exploration with Feedback Graphs.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Pure Exploration with Feedback Graphs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.626153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.626153Z digest=sha256:eeb82d19aa59b70ff0eda30e1ce5e86cc261030ae93e3fc391893d8bfdbe5b5e

Observation 2db139a3-37fb-4422-b44f-c4ca9c94165a · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.721970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.721970Z digest=sha256:cf021f24ed96980b28b4ecbc7a30862a196a54e38194ee52c6dd0948115b623e

Observation 50d7c8a2-7619-4abe-9d3d-2db699ddff46 · outbound

This paper cites Asymptotically Optimal Sequential Testing with Markovian Data.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Asymptotically Optimal Sequential Testing with Markovian Data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.796549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.796549Z digest=sha256:65f46ead01a867fcccfb49d0f01dd3fec8da85f18d027ac31903e02bc29a8ed7

Observation c8a9d37c-bec9-4521-9c69-4ffcc2cfc9e0 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.837113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.837113Z digest=sha256:23b27ed1a2cc6cd6efe8ddc37738ecbdf90c3490b8e5252e3438a6a1c5b55075

Observation a29c040a-498a-4498-9d0d-dd22925ba511 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.883532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.883532Z digest=sha256:897937059e91074c824af6f09f4496d98c943dad2de49a163d7e381f0068ac28

Observation 73eadf8b-3c6d-4b0a-bfa7-1eb5171673f6 · outbound

This paper cites Taupin, Y.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Taupin, Y

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.937068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.937068Z digest=sha256:bde75a3f81793b61b24d92fd28a961f07fa2bd58dc0bc4c08ae4cdfb98b90f60

Observation 78ae6ba3-d871-475b-9c6b-9e31467066b0 · outbound

This paper cites Tuynman and R.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Tuynman and R

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:00.975367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:00.975367Z digest=sha256:50f9aaa490a5615dd27fb428fd4663cd9e151b8a812e1ade3b588334f26fa777

Observation dc1d2a00-e8be-4281-88aa-fd7bc1184742 · outbound

This paper cites Zalinescu.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Zalinescu

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.031230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.031230Z digest=sha256:5a118a003dc8c5790e8450243f729ae7437fc8c23efc1b4f84966033d3d57a3b

Observation e79e6aa5-05e1-4189-afa2-4290b4bd6dd3 · outbound

This paper cites an unresolved cited work.

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.101340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.101340Z digest=sha256:99c5dfa126118fcbe965cab3e851a4b5d2db2983b50440331b5c0436aab2d413

Observation c115e5c2-e857-4ce1-92bb-b79578246156 · outbound

This paper cites Lemma 42.Let Ω⋆(M) := ( ω∈Ω(M) inf M′∈Alt(M) X s,a ω(s, a)KL(P(s, a), P′(s, a)) = (T⋆(M))−1 ).

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning Lemma 42.Let Ω⋆(M) := ( ω∈Ω(M) inf M′∈Alt(M) X s,a ω(s, a)KL(P(s, a), P′(s, a)) = (T⋆(M))−1 )

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:01.147324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:54:01.147324Z digest=sha256:1fe66960a3669d5ff6f48a6797b6e880a0a9b65732f387965aa6ad5818234708

Pith citing papers

No inbound Pith citation observations are available.