Pith. sign in

Paper Citation Record · LEDGER

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06521 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:10:52.644951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:45:40.553981Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 752a21c9-eded-434d-be7b-0f168f5d613f · outbound

This paper cites Navigating to the best policy in markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Navigating to the best policy in markov decision processes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.062037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.645313Z digest=sha256:3c498c02b59c597b9e0fe5768ca055046637ffabef52e48e333872f402076a20

Observation 3e9af9e9-27e2-47d8-8e8b-e7c0ff6e7545 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.650637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.650637Z digest=sha256:27d997a19dde85687d8be356cf99682f884970167719ea69962c7a35384dc2bc

Observation 2f7ec633-71fd-414e-bb52-31eb330fa6ce · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Finite-time analysis of the multiarmed bandit problem

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.022547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.655458Z digest=sha256:f699e13fefd3fe24effe97c2bd1db3b6c9379ab531926ba1068bef13137cbc0f

Observation 1560998c-4618-4634-a281-39ba3a04691c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.660747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.660747Z digest=sha256:ad299ac5b88e9404412ee680bc1fd313ff639cdef5c8c457322ded7b525c34e8

Observation e5655449-346d-46ee-9f2c-2fe69a778fac · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Minimax regret bounds for reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.665444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.665444Z digest=sha256:78a491477bc57069faaaf708b9fd87f02ba1777857310d901e7935070283ef04

Observation e54fec24-9f68-4cf5-aa49-5f8fe58a873c · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.670031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.670031Z digest=sha256:0a4cee27f9e30eeec2b9e9df9e2d76a11dfb5986b1b76d1af8871f11ecdec47b

Observation c29630a4-0f78-40e9-887c-b334b8b794a4 · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.675666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.675666Z digest=sha256:3e4595f2acbb8cc7036ca4b5a863172cf2a6810b4a07c20c767fdbf67826dffe

Observation 8f3663fe-4d78-4094-b5ae-d7cf90cf1a0d · outbound

This paper cites Top-k off-policy correction for a reinforce recommender system.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Top-k off-policy correction for a reinforce recommender system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.680902Z digest=sha256:8340e1d969d02d39783af50c84a81468f6a811a5a47ad870dc7ac089b35013cc

Observation d1ac95b1-9f22-4f96-ae91-10a8d6604ec1 · outbound

This paper cites Variance-Aware Sparse Linear Bandits.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-Aware Sparse Linear Bandits

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.085200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.685486Z digest=sha256:3c9d3d269516146156aa197c98dcacbbf2cfae0ede29bb0c6a1f4552b19cce9e

Observation cec2c58e-e241-40a0-9c94-fdd2d08dfd2b · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Policy certificates: Towards accountable reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.912679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.690575Z digest=sha256:87dc30224aee98b5a549d7fcbd740153cae0748a8e92c1989c289fec1e8ec69c

Observation 6fa8dfe2-4779-4705-9abb-c97da5ade988 · outbound

This paper cites Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.695323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.695323Z digest=sha256:6e138bb5801e7ae88f7db7e3ca2fba029615ec2641f28982eb8599d24ce7d9f0

Observation 60b87a59-dc4d-49ed-a24e-ebb846250806 · outbound

This paper cites Gap-dependent bounds for two-player markov games.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-dependent bounds for two-player markov games

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.865549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.700715Z digest=sha256:5e35eccdb53f39f7a8b3d9c81235b81b5e78e1391da8b7082b188be8d748e497

Observation a168d2be-fb72-4c2d-a587-6dbfc37b7b0a · outbound

This paper cites Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.842175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.705363Z digest=sha256:663575e897b2e9f19132a5ec02fe4931e99cc6fc9b3d33e8b8e1d5e4262642fe

Observation d22e8a7b-3d5e-480c-8a43-fe6b483cb5ad · outbound

This paper cites Efficient bias-span-constrained exploration-exploitation in reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Efficient bias-span-constrained exploration-exploitation in reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.820123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.709935Z digest=sha256:edb0f11ae97ef1d1ed5336d2529a6851d9f29f1fb5ecb200296852781edc161b

Observation ec4381ac-2eb4-48cf-9940-8968b0f5725d · outbound

This paper cites Logarithmic regret for reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic regret for reinforcement learning with linear function approximation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.800170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.714557Z digest=sha256:8ee3367fdd7606a8bca33a1e5a525728840b04ad36b9bc7a565c03420a752c94

Observation 0076f097-b362-4663-b39d-45d547528fa1 · outbound

This paper cites Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.780075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.719607Z digest=sha256:7db4c01e03beb4dd1c4c9bf912b718eedd218767535bb23a13ec8d8f5029f4e7

Observation 4c146e2b-e269-42fb-9c88-5f62ed713de8 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.723948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.723948Z digest=sha256:8dcc27262053e318d4f4e0b4d30436946575491afa254e9dc9a46cc2f3a699ee

Observation aaf290a2-8bd0-4d6a-927c-49934c46ac58 · outbound

This paper cites Reward-free exploration for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reward-free exploration for reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.751095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.728273Z digest=sha256:9a508496e1b647de0d41a8af83faa83c30885e48ef649551d9528f633871ed02

Observation 4001edbd-21f0-4c19-b55c-fc39d77c45e5 · outbound

This paper cites Planning in markov decision processes with gap-dependent sample complexity.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Planning in markov decision processes with gap-dependent sample complexity

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.729062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.732692Z digest=sha256:e7eb7e90fa6f8cabe5ba430c5d97012cd82ea69c4a907db5cea02793933f46d3

Observation e9893598-862c-4221-92c3-14e79de722bd · outbound

This paper cites Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.708304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.737357Z digest=sha256:56139645425d1a51190e54eae3b2c75153e25d6aec65556167f14161869b4851

Observation e3b9b6b2-9766-4c14-9ef0-847e3dce72d3 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Asymptotically efficient adaptive allocation rules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.742030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.742030Z digest=sha256:f24b3f7cc4d21b1d572c42d184881256298fc287ac752ad47b6df067618f23e5

Observation e475c0b6-34e8-4d90-b0ca-53a9292c24d2 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.675579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.746775Z digest=sha256:80de89e2d5ce2e508feccc199f9444d9acc0ee304ebc44ec26b4d88a081e66cc

Observation 0374572d-fe69-4408-9c01-0759e382bb4f · outbound

This paper cites Continuous control with deep reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Continuous control with deep reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.751752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.751752Z digest=sha256:9249c1b3f1e729ba7774436c699d39aedfa43cfc8329f7b56133713085c95c37

Observation d0a8208e-7d51-4449-b65d-cd43cbc729b9 · outbound

This paper cites Deep reinforcement learning for dynamic treatment regimes on medical registry data.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Deep reinforcement learning for dynamic treatment regimes on medical registry data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.655567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.756482Z digest=sha256:b04c6fb3fb1559815fc164d36b5b3e8e247f9e937c68e7501ccf311bf21b4cb0

Observation dfa9340b-fa18-408b-bcbb-1d0b634b2f37 · outbound

This paper cites Adaptive Sampling for Best Policy Identification in Markov Decision Processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Adaptive Sampling for Best Policy Identification in Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.039979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.761176Z digest=sha256:584e536b499f9872a68d879e74a173c2c0ee9b6fac0d75eada012aa1d49169af

Observation 82315765-0338-43ad-bf3a-22ed3b6f355b · outbound

This paper cites Empirical Bernstein Bounds and Sample Variance Penalization.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Empirical Bernstein Bounds and Sample Variance Penalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.766080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.766080Z digest=sha256:fd056c3033e114596489847bdda9f253533edfd28918c50befeff2a6642878d5

Observation da871a9a-82bc-4c25-9641-0600ed1a9997 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Ucb momentum q-learning: Correcting the bias without forgetting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.634394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.771025Z digest=sha256:e07e67ad10a24742beb2afb3338484c5cced063c244e4b87fb049f9f788ee86b

Observation 7e4070e2-adca-46e5-9b69-c550d4538112 · outbound

This paper cites Reinforcement learning for optimized trade execution.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning for optimized trade execution

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.612325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.775540Z digest=sha256:5475c6e58bb1eb5d13cb281a0c11f830df6fe6393a9608f3e8a2cbf0d28af38b

Observation 840d6741-43ff-467c-9513-0fb223349c7e · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.592062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.779948Z digest=sha256:7ab166675fa8e18c6a4602b56bf7ad6c7571c9cbc2f9feceef4482ef39cf89e6

Observation 83035fc7-3068-41b5-9a09-6380cec19577 · outbound

This paper cites Exploration in structured reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Exploration in structured reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.572522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.784415Z digest=sha256:04f1490a6737b00d450ebe1c325e19c8c85559950bb6c027064a7ff8ddc64e30

Observation 143457cc-037a-4666-9f5a-426af5f8d651 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.552330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.788946Z digest=sha256:a36f99937be556dc6ea4e777aec02d2328683799bd982f8f5937fbc5eb6a10c7

Observation dc82e5fe-f4e1-49a8-9b08-93ef8f346b0c · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.530992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.793430Z digest=sha256:4184e734e87605aec3d7b2ea211d687046224032d68b7238da7b9f70de1161d5

Observation fdcbba15-75cf-4a2c-8f4e-d664249465b1 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Mastering the game of go with deep neural networks and tree search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.797839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.797839Z digest=sha256:f4e16310dccf349e6636da52a5618bf32c7cac9ba869830a84e4d7827fba0a7c

Observation 1c0afb3a-986d-4ccd-8bb4-e9c44f172cfb · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Non-asymptotic gap-dependent regret bounds for tabular mdps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.802739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.802739Z digest=sha256:8a43f4d71da505196749ae726366c35adb98b88478d44f0cc772c38e4a02f1db

Observation 53f03b79-cce4-4b1a-a750-b79a831414fd · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.806964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.806964Z digest=sha256:a7dd417909275d82c2aa1dee39d7af61f65722e48f008310dccd733025625794

Observation ce9bbdc6-dc6b-45b5-b181-dd2192b8093f · outbound

This paper cites Variance-aware regret bounds for undiscounted reinforcement learning in mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-aware regret bounds for undiscounted reinforcement learning in mdps

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.481576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.811736Z digest=sha256:5d4d9361a48f21d4cf96676822cb4ed8283274618e69121787e88a442410bda9

Observation a7757793-65e0-4df2-a5bc-6aca707c4cd1 · outbound

This paper cites Optimistic linear programming gives logarithmic regret for irreducible mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic linear programming gives logarithmic regret for irreducible mdps

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.460484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.816074Z digest=sha256:a3605a93c389a68e807015a4cdea97797a1b8f40640745aaff45608ecb02abc3

Observation 87b74cce-14e2-47e2-a05b-8814cb258c4c · outbound

This paper cites Near instance-optimal pac reinforcement learning for deterministic mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near instance-optimal pac reinforcement learning for deterministic mdps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.441791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.820781Z digest=sha256:bb27a63fdf7a0eff7836a8499e9b5545e56f28787b80705715b2937a83f96e29

Observation 646c21e6-60f6-4282-ba5a-dbe02e9614a8 · outbound

This paper cites Optimistic pac reinforcement learning: the instance-dependent view.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic pac reinforcement learning: the instance-dependent view

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.422810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.825041Z digest=sha256:29fc751498f906f542c124dd7a5d8d273b7dc60acbf58088a1711249f1600eb0

Observation 665fbd0f-d338-43cc-89fe-2f62a84c2986 · outbound

This paper cites Reinforcement learning with logarithmic regret and policy switches.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning with logarithmic regret and policy switches

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.403477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.830888Z digest=sha256:1addb2fb6e2cbd6c12f954e9d9ce070644d12631786fe23c6c5204f60f1ec385

Observation a4fba1ab-3745-4ef1-9a60-f3f745eabfc8 · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.387794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.835885Z digest=sha256:a9240da751ef8058b358e83f4249a21c20bc3672a1852bc219de7bcc108a8e34

Observation c96a5531-24de-4071-ac6b-bb1c7376ee50 · outbound

This paper cites First-order regret in reinforcement learning with linear function approximation: A robust estimation approach.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs First-order regret in reinforcement learning with linear function approximation: A robust estimation approach

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.371017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.841054Z digest=sha256:4f23ec49b830281c4eb4ab65486ccdb4d9ab69420000d0e5da70e091d398128d

Observation 343a71d3-ab56-4d4c-a37e-13f8cab7b505 · outbound

This paper cites Beyond no regret: Instance-dependent pac reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond no regret: Instance-dependent pac reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.352868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.846110Z digest=sha256:9d02f2b9b4499c10f892c33ff61de2aa63084e465bbba958d4012226abc1deb7

Observation 587e52c9-d96d-4887-80a9-2f0250ffdba0 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On gap-dependent bounds for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.335760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.851766Z digest=sha256:149fae33a51eb674bb42fe281821520ef1e17da2452476df77a6b026d3d908e9

Observation 0c2295a7-13c1-4934-a9f2-7106bdeb9912 · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal randomized exploration for tabular markov decision processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.318255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.857671Z digest=sha256:fcea123ba99077608385d12afd2fc00615a31be12f1d004317905c2c3cf5ba32

Observation ee3fccd1-e11e-4db4-9784-a506b0ca06b9 · outbound

This paper cites Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.300451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.863405Z digest=sha256:e0293a76420d8e16bc9b3d1d2f9d163c3e1a15161115d444660a9ddd33f6d45d

Observation c561d791-c61e-41ee-922e-0c6d2e04dd84 · outbound

This paper cites Q-learning with logarithmic regret.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Q-learning with logarithmic regret

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.279268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.868774Z digest=sha256:e739203e50e1cda7f3a718d1aff93343aea5e32edeeac1045563c10236ab428b

Observation e09e3872-2afb-4001-84da-90d806ce96b6 · outbound

This paper cites Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.875678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.875678Z digest=sha256:517a7ac4c0c56b00825eb4a2f44f47f14601645a24b5f9b5f9320120aa09c5f6

Observation 1fdf739b-4ba0-4d61-acb9-a387c1a54368 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.246652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.882653Z digest=sha256:5b01733e03a11cfb2828e11aa6027394c242a3eb371a94ff0fb340de04d4bf87

Observation 03367650-80df-4da6-97d3-4005db1289cb · outbound

This paper cites Almost optimal model-free reinforcement learningvia reference-advantage decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Almost optimal model-free reinforcement learningvia reference-advantage decomposition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.224318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.889570Z digest=sha256:a274f8e8f80d5212ef8aafa78a0fe8d483f04c951e44dd81b07c36143019ffa3

Observation a62ee762-7be0-4634-a9f7-806d703c4041 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.203614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.897012Z digest=sha256:db43ca29fd5d654cc145d9bc65ccb19f2cfd8d250cc7c1d92906c70997cade43

Observation eb10c7a5-ddd7-49d9-9810-6017f98d517a · outbound

This paper cites Improved variance-aware confidence sets for linear bandits and linear mixture mdp.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved variance-aware confidence sets for linear bandits and linear mixture mdp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.184382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.903919Z digest=sha256:f32860b487c026c6ea328cdee0ea1e6ea211f262e362ff4791fe582030a2f76d

Observation 437ff0ee-9e1e-4079-b4c7-1dd3ac6a4efa · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.165150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.911067Z digest=sha256:9913a4095732c0374e9704c9621362d388eac8c661ae8e021697797b8280f7fa

Observation 744f8b4a-6ecc-44b7-bd4a-b4cd2893f646 · outbound

This paper cites Settling the sample complexity of online reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Settling the sample complexity of online reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.917081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.917081Z digest=sha256:a5e5715ebb415ad0d4e1f0a93d3e5424904f97a98fcb31c08d77c2c9191d46fd

Observation b1f3bd36-1d48-4a40-8458-ac15aa3701e9 · outbound

This paper cites Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:21.992642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.923994Z digest=sha256:bbeedaca5d5c11faf9a92a970c00b246173ca50672cc8764b3f6a216c871ae74

Observation a74689ed-21ec-46a7-a988-4e5560f21227 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.133022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.931913Z digest=sha256:ff7e85a4e750b91b6d82b5d4c302eb9776ee4f8577d44136c871e478bed1a54e

Observation 46abb972-bf09-441b-a206-dd59d2d18436 · outbound

This paper cites Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.939768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.939768Z digest=sha256:748ec2de521306babac928bdd2131f727e55df8fbb77eb618dc5bce7fc067b62

Pith citing papers

Observation 229aedfe-0a81-41e4-9f90-efbe3fc4d041 · inbound

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms cites this paper.

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:40.556056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:10:52.644951Z digest=sha256:213f1a97e8ed29215f2fe3c77fb442e177b813c20afe41edf16e3b3f4b2b1ef4