Pith. sign in

Paper Citation Record · LEDGER

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.06521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06521 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:10:52.644951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:45:40.553981Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy38
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 752a21c9-eded-434d-be7b-0f168f5d613f · outbound

This paper cites Navigating to the best policy in markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Navigating to the best policy in markov decision processes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.062037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.645313Z digest=sha256:5f0c1fe96ae66f44f1124709e61d6fc1b9c699270ad02ed952c54c70609d31c2

Observation 3e9af9e9-27e2-47d8-8e8b-e7c0ff6e7545 · outbound

This paper cites Logarithmic online regret bounds for undiscounted reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic online regret bounds for undiscounted reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.650637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.650637Z digest=sha256:1bb48a56ee01ec1717e6d6666d019145dd93dc7f704adddc9a0761d502ada5eb

Observation 2f7ec633-71fd-414e-bb52-31eb330fa6ce · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Finite-time analysis of the multiarmed bandit problem

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:23.022547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.655458Z digest=sha256:341890660aff48f55dc4aa17910d93840811969241734a8d264858b0723de7d1

Observation 1560998c-4618-4634-a281-39ba3a04691c · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal regret bounds for reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.660747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.660747Z digest=sha256:8c8f88b43d3854e127a1b8696930dba6d3a2b4bdcc10b48b37e8335a105d5c5f

Observation e5655449-346d-46ee-9f2c-2fe69a778fac · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Minimax regret bounds for reinforcement learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.665444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.665444Z digest=sha256:a63624484af5c17a7df3468db77a948bfbfa8d57c2fd467ff3fa96cc86d2fd55

Observation e54fec24-9f68-4cf5-aa49-5f8fe58a873c · outbound

This paper cites REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.670031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.670031Z digest=sha256:06a37bc5f0e21b2da614d257f5765c0a21bbd2c341775fa9db589b27daf33bc2

Observation c29630a4-0f78-40e9-887c-b334b8b794a4 · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.675666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.675666Z digest=sha256:3dc06ffd67118c7c11156ddcdfaf1ef556cc6bf52601590e7de66597269f99df

Observation 8f3663fe-4d78-4094-b5ae-d7cf90cf1a0d · outbound

This paper cites Top-k off-policy correction for a reinforce recommender system.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Top-k off-policy correction for a reinforce recommender system

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.680902Z digest=sha256:a1e63fdc92e1f41a06d59ff2a5239a866a186c83fe581ad0c4b09593b4b9df43

Observation d1ac95b1-9f22-4f96-ae91-10a8d6604ec1 · outbound

This paper cites Variance-Aware Sparse Linear Bandits.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-Aware Sparse Linear Bandits

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.085200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.685486Z digest=sha256:eeef6ebe7751febed43c4b9a04ab7df43af48f68e2386dfe7d2b4e533d1a600d

Observation cec2c58e-e241-40a0-9c94-fdd2d08dfd2b · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Policy certificates: Towards accountable reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.912679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.690575Z digest=sha256:d71c41f39624425d62310380d83f6e33ea3b78d0c8707c4c9710d7dedea700ed

Observation 6fa8dfe2-4779-4705-9abb-c97da5ade988 · outbound

This paper cites Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.695323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.695323Z digest=sha256:4aa2a90d32037710940e2f072bb8353ce8b21d0bff3bdfd47d98b75e6642bd96

Observation 60b87a59-dc4d-49ed-a24e-ebb846250806 · outbound

This paper cites Gap-dependent bounds for two-player markov games.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-dependent bounds for two-player markov games

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.865549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.700715Z digest=sha256:01448fc44a417acbaa55638071416386279afcb69232a18c255e4ae427918903

Observation a168d2be-fb72-4c2d-a587-6dbfc37b7b0a · outbound

This paper cites Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.842175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.705363Z digest=sha256:80b1353ec34aee6d701697648817021bf9a94e8d368d843be0c80984c07d197e

Observation d22e8a7b-3d5e-480c-8a43-fe6b483cb5ad · outbound

This paper cites Efficient bias-span-constrained exploration-exploitation in reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Efficient bias-span-constrained exploration-exploitation in reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.820123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.709935Z digest=sha256:677855c01a23001e2a5a9d4a79da44287c5026729c715e5bd607dc3735ea124a

Observation ec4381ac-2eb4-48cf-9940-8968b0f5725d · outbound

This paper cites Logarithmic regret for reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic regret for reinforcement learning with linear function approximation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.800170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.714557Z digest=sha256:3306fb7c396c0249c1dc2a136575cc4e5512103507adb3350a531a6b61977777

Observation 0076f097-b362-4663-b39d-45d547528fa1 · outbound

This paper cites Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.780075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.719607Z digest=sha256:19a090bd4c41bffcd13621da9bb9c159750bffd8ae4d5b1e091dede3a74d9693

Observation 4c146e2b-e269-42fb-9c88-5f62ed713de8 · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.723948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.723948Z digest=sha256:356701aa7eadc0e342a17dbbe3b28d9cb672f6bae3d1b83229d06d6659b9bf31

Observation aaf290a2-8bd0-4d6a-927c-49934c46ac58 · outbound

This paper cites Reward-free exploration for reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reward-free exploration for reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.751095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.728273Z digest=sha256:201fe7aa0e4650926b585b93039572887f3c2218aed46a9a4cf1da863cad7d23

Observation 4001edbd-21f0-4c19-b55c-fc39d77c45e5 · outbound

This paper cites Planning in markov decision processes with gap-dependent sample complexity.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Planning in markov decision processes with gap-dependent sample complexity

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.729062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.732692Z digest=sha256:e0fb6892b97811c47222af86ee8e33f4149061b8a11df2c5b657a26cd7978575

Observation e9893598-862c-4221-92c3-14e79de722bd · outbound

This paper cites Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.708304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.737357Z digest=sha256:8d3f7c999f03261954fbf394036aa27f56f29355066771876238aabed7e06cee

Observation e3b9b6b2-9766-4c14-9ef0-847e3dce72d3 · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Asymptotically efficient adaptive allocation rules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.742030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.742030Z digest=sha256:c1de47215a3f7a5db3a271b7037907b94afcbdd6a196e39723c159794c249ad8

Observation e475c0b6-34e8-4d90-b0ca-53a9292c24d2 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.675579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.746775Z digest=sha256:fb8e92939c7fe0144098a816a4204755a7adcb5ea5c3cff398d3118b9f7646ba

Observation 0374572d-fe69-4408-9c01-0759e382bb4f · outbound

This paper cites Continuous control with deep reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Continuous control with deep reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.751752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.751752Z digest=sha256:862ef4f19fa9d8618ce3d2ed18cdb11ca4a5bc36e41b5c85c47bb8668f742b65

Observation d0a8208e-7d51-4449-b65d-cd43cbc729b9 · outbound

This paper cites Deep reinforcement learning for dynamic treatment regimes on medical registry data.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Deep reinforcement learning for dynamic treatment regimes on medical registry data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.655567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.756482Z digest=sha256:b7c075ed8982b723425a1c2a7a7cd010e56dc0a223a3ce3b31329b112f38df31

Observation dfa9340b-fa18-408b-bcbb-1d0b634b2f37 · outbound

This paper cites Adaptive Sampling for Best Policy Identification in Markov Decision Processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Adaptive Sampling for Best Policy Identification in Markov Decision Processes

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:22.039979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.761176Z digest=sha256:e45a2fe93389c3595dc71b1970e6419ed4405d444ad16d92fe3621ea916f5ae2

Observation 82315765-0338-43ad-bf3a-22ed3b6f355b · outbound

This paper cites Empirical Bernstein Bounds and Sample Variance Penalization.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Empirical Bernstein Bounds and Sample Variance Penalization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.766080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.766080Z digest=sha256:57de02d29b8608cfcd18c8c4b64ed76e2c19c8c2e2a2f8dc498ce829a6ede4bc

Observation da871a9a-82bc-4c25-9641-0600ed1a9997 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Ucb momentum q-learning: Correcting the bias without forgetting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.634394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.771025Z digest=sha256:5f8965a0f6932e03fe6f764b89c74b633b070cce5cd00fac24b12e3e80a55535

Observation 7e4070e2-adca-46e5-9b69-c550d4538112 · outbound

This paper cites Reinforcement learning for optimized trade execution.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning for optimized trade execution

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.612325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.775540Z digest=sha256:24ea8fc3dd66aec1a13f6ee30991d1a264527aa62134f6a6a1f04f52e1fd4357

Observation 840d6741-43ff-467c-9513-0fb223349c7e · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.592062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.779948Z digest=sha256:7c7116550352fd5029ab2e495d37b368b72d5c56116886945d35306c47ad34b2

Observation 83035fc7-3068-41b5-9a09-6380cec19577 · outbound

This paper cites Exploration in structured reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Exploration in structured reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.572522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.784415Z digest=sha256:89293d170a54b0a1ac82901679371f4b7ce8e01fc69ab3811d84ec050e2f7062

Observation 143457cc-037a-4666-9f5a-426af5f8d651 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.552330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.788946Z digest=sha256:2c862f94e1f24b6dc6aa0ef0d862e4eace9e25a1ca878ed401fd73678d118530

Observation dc82e5fe-f4e1-49a8-9b08-93ef8f346b0c · outbound

This paper cites Reinforcement learning in linear mdps: Constant regret and representation selection.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning in linear mdps: Constant regret and representation selection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.530992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.793430Z digest=sha256:350dad3ce067308d0dd494a7974d9dabf716c6da500b7007e3bdf093040db7fc

Observation fdcbba15-75cf-4a2c-8f4e-d664249465b1 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Mastering the game of go with deep neural networks and tree search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.797839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.797839Z digest=sha256:d03bf44f7a8018f10deaa1e24209e230bffe49f9241fe568aebcfa64ad0f1c9f

Observation 1c0afb3a-986d-4ccd-8bb4-e9c44f172cfb · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Non-asymptotic gap-dependent regret bounds for tabular mdps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.802739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.802739Z digest=sha256:0ad46c75313c7e13dd34b01a52ae53f2e80821e85db70e11cfddf58df68f2f26

Observation 53f03b79-cce4-4b1a-a750-b79a831414fd · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning: An introduction, volume 1

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.806964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.806964Z digest=sha256:3771bf6dba362b85e3146883f5f61374f951b62f23d100cd9f55eb5cff6be7c1

Observation ce9bbdc6-dc6b-45b5-b181-dd2192b8093f · outbound

This paper cites Variance-aware regret bounds for undiscounted reinforcement learning in mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-aware regret bounds for undiscounted reinforcement learning in mdps

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.481576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.811736Z digest=sha256:50e20af9a8357ed7f3d56c56bf4efbeada6e9d50c364856c27a4fbd78aa41cf6

Observation a7757793-65e0-4df2-a5bc-6aca707c4cd1 · outbound

This paper cites Optimistic linear programming gives logarithmic regret for irreducible mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic linear programming gives logarithmic regret for irreducible mdps

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.460484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.816074Z digest=sha256:264e363aae63c068036d56401e4d1ec84331f817d7164e5b0641df14159cf242

Observation 87b74cce-14e2-47e2-a05b-8814cb258c4c · outbound

This paper cites Near instance-optimal pac reinforcement learning for deterministic mdps.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near instance-optimal pac reinforcement learning for deterministic mdps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.441791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.820781Z digest=sha256:a8cdbd39e8244e5f1d48b0633fc7c4fe0d9142fc97c1c81bb9919b8a1237d9f2

Observation 646c21e6-60f6-4282-ba5a-dbe02e9614a8 · outbound

This paper cites Optimistic pac reinforcement learning: the instance-dependent view.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic pac reinforcement learning: the instance-dependent view

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.422810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.825041Z digest=sha256:2eb3a385fb55d0dc7f91e036ede948b31e8382d83742caeccdb84f23657de396

Observation 665fbd0f-d338-43cc-89fe-2f62a84c2986 · outbound

This paper cites Reinforcement learning with logarithmic regret and policy switches.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning with logarithmic regret and policy switches

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.403477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.830888Z digest=sha256:a028077674e7d33c3662b8476765862fd70823606e0557704eac286a0a118bf6

Observation a4fba1ab-3745-4ef1-9a60-f3f745eabfc8 · outbound

This paper cites Instance-dependent near-optimal policy identification in linear mdps via online experiment design.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Instance-dependent near-optimal policy identification in linear mdps via online experiment design

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.387794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.835885Z digest=sha256:e7e69664f5d513620c06138f02fbed727ee8880b170d12dce1ad9668defc19d1

Observation c96a5531-24de-4071-ac6b-bb1c7376ee50 · outbound

This paper cites First-order regret in reinforcement learning with linear function approximation: A robust estimation approach.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs First-order regret in reinforcement learning with linear function approximation: A robust estimation approach

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.371017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.841054Z digest=sha256:9475dd30de534de46b5d40e1475ec007b2c5e1f815e00c3bf9992873cffafe4c

Observation 343a71d3-ab56-4d4c-a37e-13f8cab7b505 · outbound

This paper cites Beyond no regret: Instance-dependent pac reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond no regret: Instance-dependent pac reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.352868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.846110Z digest=sha256:26999e5b7cd3776c3936853939f948f9f3a91d29d8c889d2ff4cce2e6e2a5729

Observation 587e52c9-d96d-4887-80a9-2f0250ffdba0 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On gap-dependent bounds for offline reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.335760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.851766Z digest=sha256:b002cef32f3eabd7790eebde4f01189b20472e943709b413745fd69d297f2a45

Observation 0c2295a7-13c1-4934-a9f2-7106bdeb9912 · outbound

This paper cites Near-optimal randomized exploration for tabular markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal randomized exploration for tabular markov decision processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.318255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.857671Z digest=sha256:a2f233ad33f7f4969f47879bbe81ba40459da79a372b031e87518c8d75aa91fe

Observation ee3fccd1-e11e-4db4-9784-a506b0ca06b9 · outbound

This paper cites Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.300451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.863405Z digest=sha256:822e55677812a72c4b6300847f5f0943a9ab9abb07540f905b0f1bedeb2e0b07

Observation c561d791-c61e-41ee-922e-0c6d2e04dd84 · outbound

This paper cites Q-learning with logarithmic regret.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Q-learning with logarithmic regret

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.279268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.868774Z digest=sha256:26515921f8fc2c38cadd7f56ec353d84fb683a29027494fee1d056cd6b53e497

Observation e09e3872-2afb-4001-84da-90d806ce96b6 · outbound

This paper cites Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.875678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.875678Z digest=sha256:a0f41b53b3fd625f12be2542ff463952f64f9414f4492a57a79711d62bc09b27

Observation 1fdf739b-4ba0-4d61-acb9-a387c1a54368 · outbound

This paper cites Regret minimization for reinforcement learning by evaluating the optimal bias function.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.246652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.882653Z digest=sha256:b7e7e9401148299d772e759466eec50a461b5ce9b378a1f867048c09e4a64edc

Observation 03367650-80df-4da6-97d3-4005db1289cb · outbound

This paper cites Almost optimal model-free reinforcement learningvia reference-advantage decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Almost optimal model-free reinforcement learningvia reference-advantage decomposition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.224318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.889570Z digest=sha256:e406e48587e9b1104e265ffcde5fea017d107755e262e5325e8ae3bfa06ee0e4

Observation a62ee762-7be0-4634-a9f7-806d703c4041 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.203614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.897012Z digest=sha256:99fa8784d4a50922aa68fa857f52a4234aab7e55988b2e2930c0778c8fa269b1

Observation eb10c7a5-ddd7-49d9-9810-6017f98d517a · outbound

This paper cites Improved variance-aware confidence sets for linear bandits and linear mixture mdp.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved variance-aware confidence sets for linear bandits and linear mixture mdp

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.184382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.903919Z digest=sha256:2855c8be4ee6183f3dc1d481b7d8941babaa15e4a56d72673e7e3f96960d3f91

Observation 437ff0ee-9e1e-4079-b4c7-1dd3ac6a4efa · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.165150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.911067Z digest=sha256:7dab60b741eb29307e1f8b966cab13967c423fc61e9e760a019b85114c48df62

Observation 744f8b4a-6ecc-44b7-bd4a-b4cd2893f646 · outbound

This paper cites Settling the sample complexity of online reinforcement learning.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Settling the sample complexity of online reinforcement learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.917081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.917081Z digest=sha256:3a7f2a0a370f118a1e2f4ab30e3c50c0c48758823db88bbc20c62e751b0e378d

Observation b1f3bd36-1d48-4a40-8458-ac15aa3701e9 · outbound

This paper cites Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:09:21.992642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.923994Z digest=sha256:0a27dd402fb40c0515a54c748312ed74881a00148f612bad00a453d3685ddff2

Observation a74689ed-21ec-46a7-a988-4e5560f21227 · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:09:22.133022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:09:21.931913Z digest=sha256:3ab3720cccd20b2c48f9b90359a21fac8cde4aae5c2b0f7e047ee8e0f5ccf505

Observation 46abb972-bf09-441b-a206-dd59d2d18436 · outbound

This paper cites Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments.

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:21.939768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:21.939768Z digest=sha256:a70e026af1de5cb80806ae85b2d978307ab9e8c36e44603241716cf61cf9ea6a

Pith citing papers

Observation 229aedfe-0a81-41e4-9f90-efbe3fc4d041 · inbound

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms cites this paper.

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:40.556056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:10:52.644951Z digest=sha256:a6e1801ed6779219f372b29f57c54c8e7981a61f0499afdd1bdd3ebe104edc23