Pith. sign in

Paper Citation Record · LEDGER

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06521 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:10:52.644951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:45:40.553981Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 752a21c9-eded-434d-be7b-0f168f5d613f · outbound

This paper cites Navigating to the best policy in markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Navigating to the best policy in markov decision processes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.062037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.645313Z digest=sha256:444b0506795395f719f5dbcbf923fbb3242850cdb84533d8ed198c4dfa7a2753

Observation 3e9af9e9-27e2-47d8-8e8b-e7c0ff6e7545 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.650637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.650637Z digest=sha256:1bb48a56ee01ec1717e6d6666d019145dd93dc7f704adddc9a0761d502ada5eb

Observation 2f7ec633-71fd-414e-bb52-31eb330fa6ce · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Finite-time analysis of the multiarmed bandit problem

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.022547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.655458Z digest=sha256:21d2568eed34ad9166d5a65107c07956699fdd717d4ec4f2bea438ca3ea479d3

Observation 1560998c-4618-4634-a281-39ba3a04691c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.660747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.660747Z digest=sha256:8c8f88b43d3854e127a1b8696930dba6d3a2b4bdcc10b48b37e8335a105d5c5f

Observation e5655449-346d-46ee-9f2c-2fe69a778fac · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Minimax regret bounds for reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.665444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.665444Z digest=sha256:a63624484af5c17a7df3468db77a948bfbfa8d57c2fd467ff3fa96cc86d2fd55

Observation e54fec24-9f68-4cf5-aa49-5f8fe58a873c · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.670031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.670031Z digest=sha256:06a37bc5f0e21b2da614d257f5765c0a21bbd2c341775fa9db589b27daf33bc2

Observation c29630a4-0f78-40e9-887c-b334b8b794a4 · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.675666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.675666Z digest=sha256:3dc06ffd67118c7c11156ddcdfaf1ef556cc6bf52601590e7de66597269f99df

Observation 8f3663fe-4d78-4094-b5ae-d7cf90cf1a0d · outbound

This paper cites Top-k off-policy correction for a reinforce recommender system.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Top-k off-policy correction for a reinforce recommender system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.680902Z digest=sha256:0b02e7132b0de2f93f65580c60f6f52e549517e5c1dbae183c4637663f69f596

Observation d1ac95b1-9f22-4f96-ae91-10a8d6604ec1 · outbound

This paper cites Variance-Aware Sparse Linear Bandits.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-Aware Sparse Linear Bandits

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.085200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.685486Z digest=sha256:5e60ca2fa584bbaa0d976015f27ac87d64e82a5367c75d5adf04dde184e4f452

Observation cec2c58e-e241-40a0-9c94-fdd2d08dfd2b · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Policy certificates: Towards accountable reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.912679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.690575Z digest=sha256:cf1060ee146edddadda8d4c29b55d17a3ecb40d34f4179af8e75fc63873eaa21

Observation 6fa8dfe2-4779-4705-9abb-c97da5ade988 · outbound

This paper cites Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.695323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.695323Z digest=sha256:4aa2a90d32037710940e2f072bb8353ce8b21d0bff3bdfd47d98b75e6642bd96

Observation 60b87a59-dc4d-49ed-a24e-ebb846250806 · outbound

This paper cites Gap-dependent bounds for two-player markov games.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-dependent bounds for two-player markov games

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.865549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.700715Z digest=sha256:be2892d99c89be61e1b97b23e924b7e08efde00a79f97925ab42f830b35be391

Observation a168d2be-fb72-4c2d-a587-6dbfc37b7b0a · outbound

This paper cites Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.842175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.705363Z digest=sha256:f07c88ed2148a47d7d77b9799e5f969779ebd4ea1e51534d643967d4c692d8cb

Observation d22e8a7b-3d5e-480c-8a43-fe6b483cb5ad · outbound

This paper cites Efficient bias-span-constrained exploration-exploitation in reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Efficient bias-span-constrained exploration-exploitation in reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.820123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.709935Z digest=sha256:7f045b3ed3948023f386785099f4f78298100704ac25c1843445019dd6afda9e

Observation ec4381ac-2eb4-48cf-9940-8968b0f5725d · outbound

This paper cites Logarithmic regret for reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic regret for reinforcement learning with linear function approximation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.800170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.714557Z digest=sha256:7a7d0fdf0b3dfcab24bb34f1a2800a3682ca950c95666217bd701264cb83d352

Observation 0076f097-b362-4663-b39d-45d547528fa1 · outbound

This paper cites Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.780075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.719607Z digest=sha256:2f5136d6e62681fab53a41bbb8271d7233ba0061c9ef2ecdc67ee6afad43499f

Observation 4c146e2b-e269-42fb-9c88-5f62ed713de8 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.723948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.723948Z digest=sha256:356701aa7eadc0e342a17dbbe3b28d9cb672f6bae3d1b83229d06d6659b9bf31

Observation aaf290a2-8bd0-4d6a-927c-49934c46ac58 · outbound

This paper cites Reward-free exploration for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reward-free exploration for reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.751095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.728273Z digest=sha256:26bd144851846ee896a56c4a95e374524d884742625fa2a6945ed069560af490

Observation 4001edbd-21f0-4c19-b55c-fc39d77c45e5 · outbound

This paper cites Planning in markov decision processes with gap-dependent sample complexity.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Planning in markov decision processes with gap-dependent sample complexity

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.729062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.732692Z digest=sha256:4d49839c579adfca9a52dd68cad169ae9fa485f7b0a84975ec9cc520c38112a7

Observation e9893598-862c-4221-92c3-14e79de722bd · outbound

This paper cites Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.708304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.737357Z digest=sha256:ea133926bce7d47afdc66f3141f17e4104ec26c0d389b850b268b3db0f74eab4

Observation e3b9b6b2-9766-4c14-9ef0-847e3dce72d3 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Asymptotically efficient adaptive allocation rules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.742030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.742030Z digest=sha256:c1de47215a3f7a5db3a271b7037907b94afcbdd6a196e39723c159794c249ad8

Observation e475c0b6-34e8-4d90-b0ca-53a9292c24d2 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.675579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.746775Z digest=sha256:5fb748047ab35b5c561f6ad8567a53616b90fb821616ba093ecd56f5d9a876fc

Observation 0374572d-fe69-4408-9c01-0759e382bb4f · outbound

This paper cites Continuous control with deep reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Continuous control with deep reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.751752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.751752Z digest=sha256:862ef4f19fa9d8618ce3d2ed18cdb11ca4a5bc36e41b5c85c47bb8668f742b65

Observation d0a8208e-7d51-4449-b65d-cd43cbc729b9 · outbound

This paper cites Deep reinforcement learning for dynamic treatment regimes on medical registry data.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Deep reinforcement learning for dynamic treatment regimes on medical registry data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.655567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.756482Z digest=sha256:3707d6c10ed03b81c0db3ecca5dcdf2b97539c09063e409198f9fb5c8409b24b

Observation dfa9340b-fa18-408b-bcbb-1d0b634b2f37 · outbound

This paper cites Adaptive Sampling for Best Policy Identification in Markov Decision Processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Adaptive Sampling for Best Policy Identification in Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.039979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.761176Z digest=sha256:be57676f6e49d2ca511cae79505f78c845c724d4e977c1b78cdd38dcaca584de

Observation 82315765-0338-43ad-bf3a-22ed3b6f355b · outbound

This paper cites Empirical Bernstein Bounds and Sample Variance Penalization.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Empirical Bernstein Bounds and Sample Variance Penalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.766080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.766080Z digest=sha256:57de02d29b8608cfcd18c8c4b64ed76e2c19c8c2e2a2f8dc498ce829a6ede4bc

Observation da871a9a-82bc-4c25-9641-0600ed1a9997 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Ucb momentum q-learning: Correcting the bias without forgetting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.634394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.771025Z digest=sha256:a292002430955f4fea911e7a755d62d5696733902c624605983468bcfebad2aa

Observation 7e4070e2-adca-46e5-9b69-c550d4538112 · outbound

This paper cites Reinforcement learning for optimized trade execution.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning for optimized trade execution

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.612325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.775540Z digest=sha256:8165af99baf373cf5148cd3a842f971c5ee4befe86cbd43ae12a2e61c933e14e

Observation 840d6741-43ff-467c-9513-0fb223349c7e · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.592062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.779948Z digest=sha256:8282d30038902b4526b5511045013d5dbcf04255123d069172de227c87931ac3

Observation 83035fc7-3068-41b5-9a09-6380cec19577 · outbound

This paper cites Exploration in structured reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Exploration in structured reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.572522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.784415Z digest=sha256:9589b3ea740f7ff7dc5fbe16132866eb0544a6b67f710a8973099767a192f717

Observation 143457cc-037a-4666-9f5a-426af5f8d651 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.552330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.788946Z digest=sha256:ab576780942daf5afda4c425140a4a81cac0f55a2c36fe0735c3c3e249d5d82e

Observation dc82e5fe-f4e1-49a8-9b08-93ef8f346b0c · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.530992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.793430Z digest=sha256:b4e3b33a3136f949ed2178f719e691c0820b7814d88e7edf643380e437f46f90

Observation fdcbba15-75cf-4a2c-8f4e-d664249465b1 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Mastering the game of go with deep neural networks and tree search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.797839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.797839Z digest=sha256:d03bf44f7a8018f10deaa1e24209e230bffe49f9241fe568aebcfa64ad0f1c9f

Observation 1c0afb3a-986d-4ccd-8bb4-e9c44f172cfb · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Non-asymptotic gap-dependent regret bounds for tabular mdps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.802739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.802739Z digest=sha256:0ad46c75313c7e13dd34b01a52ae53f2e80821e85db70e11cfddf58df68f2f26

Observation 53f03b79-cce4-4b1a-a750-b79a831414fd · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.806964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.806964Z digest=sha256:3771bf6dba362b85e3146883f5f61374f951b62f23d100cd9f55eb5cff6be7c1

Observation ce9bbdc6-dc6b-45b5-b181-dd2192b8093f · outbound

This paper cites Variance-aware regret bounds for undiscounted reinforcement learning in mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-aware regret bounds for undiscounted reinforcement learning in mdps

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.481576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.811736Z digest=sha256:09f2aa694f840ae9d52bf28918a27307e0ad4d95cc6dae9b71ffb62cb2cde2fe

Observation a7757793-65e0-4df2-a5bc-6aca707c4cd1 · outbound

This paper cites Optimistic linear programming gives logarithmic regret for irreducible mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic linear programming gives logarithmic regret for irreducible mdps

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.460484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.816074Z digest=sha256:d4dd312b7f7e79a75868ce197806423b854dc8b843efe9e3a08b9889ac7f1337

Observation 87b74cce-14e2-47e2-a05b-8814cb258c4c · outbound

This paper cites Near instance-optimal pac reinforcement learning for deterministic mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near instance-optimal pac reinforcement learning for deterministic mdps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.441791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.820781Z digest=sha256:4c53291c900f5ab4196829e0b5b94c0d9f4ba26c33db8539b929d0abdf17ad3f

Observation 646c21e6-60f6-4282-ba5a-dbe02e9614a8 · outbound

This paper cites Optimistic pac reinforcement learning: the instance-dependent view.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic pac reinforcement learning: the instance-dependent view

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.422810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.825041Z digest=sha256:71d57cff3fe360306402d5b0006f238512c3433e4f756f9548101bf3f2c6dfde

Observation 665fbd0f-d338-43cc-89fe-2f62a84c2986 · outbound

This paper cites Reinforcement learning with logarithmic regret and policy switches.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning with logarithmic regret and policy switches

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.403477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.830888Z digest=sha256:08ca4bcb498e3f4755e6ceef0ea19d1bb8f17dc4f851e62a1e3884eb99050bd4

Observation a4fba1ab-3745-4ef1-9a60-f3f745eabfc8 · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.387794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.835885Z digest=sha256:f9d95ada150c07eaa5a8d8e7ff14d9e3dffe7193e1d3dacad6ae182398e3ef8e

Observation c96a5531-24de-4071-ac6b-bb1c7376ee50 · outbound

This paper cites First-order regret in reinforcement learning with linear function approximation: A robust estimation approach.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs First-order regret in reinforcement learning with linear function approximation: A robust estimation approach

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.371017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.841054Z digest=sha256:2404546f0bbbcadf9c81affa9ca6b02841c768b0b8854dce8633bc2b9d8ec21f

Observation 343a71d3-ab56-4d4c-a37e-13f8cab7b505 · outbound

This paper cites Beyond no regret: Instance-dependent pac reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond no regret: Instance-dependent pac reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.352868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.846110Z digest=sha256:c20856e533129f0cdfc73edaf383ea4996ea99ae54b39de41808066e5498cea7

Observation 587e52c9-d96d-4887-80a9-2f0250ffdba0 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On gap-dependent bounds for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.335760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.851766Z digest=sha256:b89ed56f63653a18abc6b9531a0a5fb75e640cc9859eb0f4112c202ce389ecae

Observation 0c2295a7-13c1-4934-a9f2-7106bdeb9912 · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal randomized exploration for tabular markov decision processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.318255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.857671Z digest=sha256:9ddc01636a7a18cc551d318fbb014f976286de1d47d9aa682df244482c314e4c

Observation ee3fccd1-e11e-4db4-9784-a506b0ca06b9 · outbound

This paper cites Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.300451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.863405Z digest=sha256:a2ee16d363c70e26350b6e109ab000bb5caa610c34b257a572e51ea7472f5715

Observation c561d791-c61e-41ee-922e-0c6d2e04dd84 · outbound

This paper cites Q-learning with logarithmic regret.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Q-learning with logarithmic regret

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.279268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.868774Z digest=sha256:f2bb99c138a7c2325d3b240f9db24095d2a2e4b0a9e1cf076a28cb1419e681f0

Observation e09e3872-2afb-4001-84da-90d806ce96b6 · outbound

This paper cites Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.875678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.875678Z digest=sha256:a0f41b53b3fd625f12be2542ff463952f64f9414f4492a57a79711d62bc09b27

Observation 1fdf739b-4ba0-4d61-acb9-a387c1a54368 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.246652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.882653Z digest=sha256:512525cb578f6e1f29b849d0b971e33825e060e50dc4730bb0e43375370a2d95

Observation 03367650-80df-4da6-97d3-4005db1289cb · outbound

This paper cites Almost optimal model-free reinforcement learningvia reference-advantage decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Almost optimal model-free reinforcement learningvia reference-advantage decomposition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.224318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.889570Z digest=sha256:fd02d61bccf7ef81f8849e3235a154e9633c2de10e3bfd3809970154a0c62216

Observation a62ee762-7be0-4634-a9f7-806d703c4041 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.203614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.897012Z digest=sha256:b4814c944edf66f4668d31aed4b8c707e3c2e31020f072948e658756db9cdf93

Observation eb10c7a5-ddd7-49d9-9810-6017f98d517a · outbound

This paper cites Improved variance-aware confidence sets for linear bandits and linear mixture mdp.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved variance-aware confidence sets for linear bandits and linear mixture mdp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.184382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.903919Z digest=sha256:07791ad0bd3422e8fe9f9190c6ca31992475dbb897afc656d06d4254c9b649f2

Observation 437ff0ee-9e1e-4079-b4c7-1dd3ac6a4efa · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.165150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.911067Z digest=sha256:ca44af2633d2b81e95adafd6d0719d352ea5b44a86be9a7a06da3ed9c4fa8cc5

Observation 744f8b4a-6ecc-44b7-bd4a-b4cd2893f646 · outbound

This paper cites Settling the sample complexity of online reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Settling the sample complexity of online reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.917081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.917081Z digest=sha256:3a7f2a0a370f118a1e2f4ab30e3c50c0c48758823db88bbc20c62e751b0e378d

Observation b1f3bd36-1d48-4a40-8458-ac15aa3701e9 · outbound

This paper cites Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:21.992642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.923994Z digest=sha256:4a0a5cbda0c3ab1eacad84fe9630fa1e05765dddae077604ee9479e9e2c10269

Observation a74689ed-21ec-46a7-a988-4e5560f21227 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.133022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.931913Z digest=sha256:b12972149d3fd1a99d185581299541e0f68c3d215dc1cae99c41660bff8d5c8f

Observation 46abb972-bf09-441b-a206-dd59d2d18436 · outbound

This paper cites Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.939768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.939768Z digest=sha256:a70e026af1de5cb80806ae85b2d978307ab9e8c36e44603241716cf61cf9ea6a

Pith citing papers

Observation 229aedfe-0a81-41e4-9f90-efbe3fc4d041 · inbound

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms cites this paper.

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:40.556056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:10:52.644951Z digest=sha256:627df2c0d65b18caa43ead8d098264b616e381e8e1e3b151192a8e2301048eaa