Pith. sign in

Paper Citation Record · LEDGER

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19854 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:27.181643Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4801e300-7b59-41b0-8bfe-f7a0fcf0f73d · outbound

This paper cites Optimistic posterior sampling for reinforcement learning: worst-case regret bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Optimistic posterior sampling for reinforcement learning: worst-case regret bounds

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.728512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.728512Z digest=sha256:d6bf7cd608abd8d8461ec4d1417448944be7ab12088e3c081ffd80029fc329a5

Observation 6c269259-52a8-433e-bf90-68ac84d15294 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Minimax regret bounds for reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.814825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.814825Z digest=sha256:8819723e8fd1d73b44a5e4e85855cb55adc4886846ffef9b1f7324dc4ffec2c8

Observation fce1e969-1212-4c88-9cc1-30a58ceec6a6 · outbound

This paper cites Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:22.931295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:22.931295Z digest=sha256:f2968bf09ac04d485d60b28e20853ad5dcf3da5bb05b9745feca4868bf052816

Observation 7ab63c10-7f08-477c-abc8-699b27b6aeab · outbound

This paper cites Brafman and Moshe Tennenholtz.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Brafman and Moshe Tennenholtz

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.010077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.010077Z digest=sha256:787c92c01bb88ebd279d28d32ec8834936c25ba2b03a9ac0cc0d4ea32d77a1fa

Observation 9174a010-34e2-4312-9c41-61633472f8a2 · outbound

This paper cites Provably Efficient Exploration in Policy Optimization.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Provably Efficient Exploration in Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.151097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.151097Z digest=sha256:86c6b8f9d45df611828c81c0915bc8c313ca2039d5c3e0883a74f69d0150c8aa

Observation e614e620-efcd-49f5-8143-c4dec2944f3a · outbound

This paper cites Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path.Advances in Neural Information Processing Systems, 34, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.271360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.271360Z digest=sha256:3732c789cb21707afd139296b2af04ac88241f17533a4e09dfb1d5f6dab7fcc4

Observation 15128eb3-93ae-4a2b-b3f8-ae7da0dcd47e · outbound

This paper cites Sample complexity of episodic fixed-horizon reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Sample complexity of episodic fixed-horizon reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.400362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.400362Z digest=sha256:f980152af579f635abcb2b4f0502f78ca171939da63a750ebfd5beb1bf93200a

Observation 00b6e90a-50bb-482a-ba23-3db55c9c7e5c · outbound

This paper cites Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.552816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.552816Z digest=sha256:d37ddd3fb31fc7a2fe351b8b30db49cef6db5cf3a6799450dd21326b04a8ba98

Observation 31b74b9f-edf2-4c00-9f0f-fa38aafa09cd · outbound

This paper cites Policy certificates: Towards accountable reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Policy certificates: Towards accountable reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.740350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.740350Z digest=sha256:5e5a33f3f480c8af371834cae09a3c58a1459bcebf01057ef54531694bf39e25

Observation 4c7bf4ac-ee8b-40aa-8ac7-de0c812c5dc2 · outbound

This paper cites Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.836348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.836348Z digest=sha256:6300c62cf82d83c7947368be0c5aa1a7371df06a1691a30a5be32a887c1f0f1b

Observation 5faac2ce-e8c0-48ab-846a-259c55210303 · outbound

This paper cites Near optimal exploration-exploitation in non- communicating markov decision processes.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near optimal exploration-exploitation in non- communicating markov decision processes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.893443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.893443Z digest=sha256:a88d3dbe39961ba0098dee8fcdabd24315ee9bcefb5419d18c48ef24f5e30aa4

Observation e095135e-2139-42b5-83c1-2f0f2157a5a9 · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal regret bounds for reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.944648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.944648Z digest=sha256:c75f68f5a8c619a7cfcd3c394fbd4a5d7ed57f4b659503a4127bd2ac21865f77

Observation 7c6d83dd-0910-48af-a758-a4812b4c8f5b · outbound

This paper cites Open problem: The dependence of sample complexity lower bounds on planning horizon.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Open problem: The dependence of sample complexity lower bounds on planning horizon

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:23.994297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:23.994297Z digest=sha256:1acc42cb6491a34ed3116ca8a30f2981c6860ec64e507b5261ab818aa77b9be1

Observation 7c7db996-bc4e-4ab8-8353-0bc018a3330a · outbound

This paper cites Is Q-learning provably efficient? InAdvances in Neural Information Processing Systems, pages 4863–4873, 2018.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is Q-learning provably efficient? InAdvances in Neural Information Processing Systems, pages 4863–4873, 2018

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.079508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.079508Z digest=sha256:d57a8f980e245395187f38237d08426ca33b48e0077cc60910a416dad2459c2c

Observation 98353e53-0d49-439a-8d18-a6ecee66b4d0 · outbound

This paper cites PhD thesis, University of London London, England, 2003.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PhD thesis, University of London London, England, 2003

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.157975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.157975Z digest=sha256:da5c3ab93063645c8a9e3a86a95976088d5f66c8d2d398b8117ab39a2d0952ed

Observation 489bb316-9a4c-4425-a6ed-eabe06adb878 · outbound

This paper cites Near-optimal reinforcement learning in polynominal time.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal reinforcement learning in polynominal time

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.268544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.268544Z digest=sha256:82a95f5baaae695d589f5bab77799b3a98c3543675609da8e091263e7e999d5d

Observation fa7e42a6-9880-4103-857a-5b5834278656 · outbound

This paper cites Near-bayesian exploration in polynomial time.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-bayesian exploration in polynomial time

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.384506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.384506Z digest=sha256:d1a97af9237ac322fed5c8fb16d9a9423daec0acd5be7a76a3c78252542e33c9

Observation d2da9399-ed90-4a4c-970b-8993b1be5470 · outbound

This paper cites Pac bounds for discounted mdps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Pac bounds for discounted mdps

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.436716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.436716Z digest=sha256:bbce41ac7716609bee6038c561e7dd0b73ab49a9a5b9f6a7ba317e622876d921

Observation 59bf1383-2265-46ad-8a9c-9bb2823128b5 · outbound

This paper cites Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.Advances in Neural Information Processing Systems, 34, 2021

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.528200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.528200Z digest=sha256:c3f81ea020ddda03cd7c05fe8ad6433efb2dde861eaf3082745c856bf6eadc51

Observation 28ca7571-374a-4447-bd9f-0c3aaceb0e04 · outbound

This paper cites Horizon-free learning for Markov decision processes and games: Stochasti- cally bounded rewards and improved bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free learning for Markov decision processes and games: Stochasti- cally bounded rewards and improved bounds

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.622410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.622410Z digest=sha256:5c287b2e3d5d8bd3573e1591aa23855b08ae8c1f7f3ceba826bec6902f4c81f3

Observation ff56f70c-8e43-440a-ada9-b1c6b75b78d2 · outbound

This paper cites Settling the horizon-dependence of sample complexity in reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Settling the horizon-dependence of sample complexity in reinforcement learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.723668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.723668Z digest=sha256:6c4c6a40a2dc2c3b4c8c4a67e17d7788d541110658e9ba782929bf4a661a2933

Observation bd057107-90c4-4b3a-a149-94e9ae10c6db · outbound

This paper cites Empirical Bernstein bounds and sample variance penalization.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Empirical Bernstein bounds and sample variance penalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.839310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.839310Z digest=sha256:7d59ca12a71e4ec33aeff8d05a50ecac3851b638968e650979d7008d9c8e6f95

Observation 03d8f37b-4764-4631-8a16-006a14384333 · outbound

This paper cites Ucb momentum q-learning: Correcting the bias without forgetting.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Ucb momentum q-learning: Correcting the bias without forgetting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:24.923691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:24.923691Z digest=sha256:11ba097d1bddcfe0f93a9ee69d95c65f710b99080b8f1ff9356900767dd7fb75

Observation 6f199159-8fd5-4a68-8c4a-4f7a24c25b5c · outbound

This paper cites A Unifying View of Optimism in Episodic Reinforcement Learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence A Unifying View of Optimism in Episodic Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.011502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.011502Z digest=sha256:0b574f57cf0adf9b594dc8b8cfad4a2d00c2208d36a8022af4fd7131230f2c5b

Observation 25829e35-b5ec-41be-90bf-7e226d8fc955 · outbound

This paper cites (more) efficient reinforcement learning via posterior sampling.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence (more) efficient reinforcement learning via posterior sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.130168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.130168Z digest=sha256:6a0883280458342a8af8ce0ac550f3a41cc8046fb81875a63a5d2bd531536694

Observation 245af5ec-f3da-421a-ba75-7cf266ec160d · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? InProceedings of the 34th International Conference on Machine Learning-Volume 70.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Why is posterior sampling better than optimism for reinforcement learning? InProceedings of the 34th International Conference on Machine Learning-Volume 70

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.306243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.306243Z digest=sha256:1a94ebc47f4ce3551f5e4fcddd8813195c0acf8cb78fd724a88c3244c6d32876

Observation b6a3e0ad-2587-48ca-ac34-ade72daa83ed · outbound

This paper cites Towards Tractable Optimism in Model-Based Reinforcement Learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Towards Tractable Optimism in Model-Based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.432600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.432600Z digest=sha256:dad341e12f6551f30f04c7bb6f4776a9d6d2da838bd48375b9bbe9aa95adbdba

Observation 245c824c-5e34-434d-898a-a73e470f401f · outbound

This paper cites Nearly horizon-free offline reinforcement learning.Advances in neural information processing systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Nearly horizon-free offline reinforcement learning.Advances in neural information processing systems, 34, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.606430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.606430Z digest=sha256:3aa3772e3141f470896d75cf29b8fceeafba4633826fc3ec7f7d97f43f3a1441

Observation d2f5be01-be32-4e1e-b9a4-f79376a246d1 · outbound

This paper cites Worst-case regret bounds for exploration via randomized value functions.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Worst-case regret bounds for exploration via randomized value functions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.723620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.723620Z digest=sha256:1b743de7e2a99ff6026a9561f64d00fb6cc1adec4da0128b7ad323bd8013c945

Observation 4e8d873f-f2c7-4b34-90fe-a89e81992529 · outbound

This paper cites Non-asymptotic gap-dependent regret bounds for tabular MDPs.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Non-asymptotic gap-dependent regret bounds for tabular MDPs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.800628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.800628Z digest=sha256:b6a30ee791a800d7d4f5438f533cdb5993248cec1a0b9a90df21f9ca9c8d3843

Observation 75b6acb7-2315-4367-8bdb-d6e9152453a7 · outbound

This paper cites PAC model-free reinforcement learning.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PAC model-free reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.889128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.889128Z digest=sha256:88a6f2d9dbf9274b8590fe8af9041f4e87f8df6516a1c06548ff785353ff7b10

Observation adc7cf24-c48e-4411-acde-1175eb5923a7 · outbound

This paper cites An analysis of model-based interval estimation for markov decision processes.Journal of Computer and System Sciences, 74(8):1309–1331, 2008.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence An analysis of model-based interval estimation for markov decision processes.Journal of Computer and System Sciences, 74(8):1309–1331, 2008

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:25.980493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:25.980493Z digest=sha256:d20f9e2cf9c5592815aacf16d79d79a1d7f7d95b8ec69d89fb5f963076add53b

Observation 59520a41-9a07-4041-af11-e2acddf13094 · outbound

This paper cites Model-based reinforcement learning with nearly tight exploration complexity bounds.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Model-based reinforcement learning with nearly tight exploration complexity bounds

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.075780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.075780Z digest=sha256:c75b8cb2ed000b68b7ef05ef7b4215f06a50e9186e82815c64123fd48958ae7e

Observation d18ebd5e-b905-4ede-a069-f6526fb93e1a · outbound

This paper cites Variance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.138252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.138252Z digest=sha256:f5c0295702ce9fed3083012fb51ec2392fd53b844c1bb55a645bfe0dee0014e2

Observation 8f2af213-6957-4163-b8e0-5c1528a40e63 · outbound

This paper cites Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret.Advances in Neural Information Processing Systems, 34, 2021.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret.Advances in Neural Information Processing Systems, 34, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.236474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.236474Z digest=sha256:04758a49c48da6818f8cf3da0224c0c2b1454c97d2aadc929cbe981e6d4aeedd

Observation d572841b-2c0e-452e-bc99-aa6a402735a8 · outbound

This paper cites Is long horizon reinforcement learning more difficult than short horizon reinforcement learning? InAdvances in Neural Information Processing Systems, 2020.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is long horizon reinforcement learning more difficult than short horizon reinforcement learning? InAdvances in Neural Information Processing Systems, 2020

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.311969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.311969Z digest=sha256:63d59a6b39b8455be5d31af9a275be18c32e9958b87c703003a8ff23f33d3e11

Observation 041767f7-ac81-47bd-a7a1-123a5d521c32 · outbound

This paper cites Near-Optimal Randomized Exploration for Tabular Markov Decision Processes.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-Optimal Randomized Exploration for Tabular Markov Decision Processes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.383570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.383570Z digest=sha256:5f55caf0a7bb9ca0b67d61fb3ce5076a2611b9599693dab328d3a70519e16c7a

Observation de55941f-e943-4f1b-875e-15a01520adff · outbound

This paper cites $Q$-learning with Logarithmic Regret.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence $Q$-learning with Logarithmic Regret

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.463961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.463961Z digest=sha256:0a0e29241ba2da205d8551e2fc0228fed3f00a6d8e82810f34674444b226f96d

Observation e0baaf33-cc64-4e9b-afe5-77e15b5ab775 · outbound

This paper cites Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.560944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.560944Z digest=sha256:81022706633b0d0ac8fd530762b88c209cac742eb2b8e461924a5c0a6b78bf6b

Observation c354c13d-6c92-473b-8046-e54aba83750e · outbound

This paper cites Horizon-free reinforcement learning in polynomial time: the power of stationary policies.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free reinforcement learning in polynomial time: the power of stationary policies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.624890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.624890Z digest=sha256:6b6c67f9b758e50a1cb80a5d0cf7abbf318be55d19f783dc5a1d74e98f923ecb

Observation 06a19116-357c-426b-8a81-1be9e439aa0f · outbound

This paper cites Variance-aware confidence set: Variance- dependent bound for linear bandits and horizon-free bound for linear mixture mdp.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-aware confidence set: Variance- dependent bound for linear bandits and horizon-free bound for linear mixture mdp

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.701935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.701935Z digest=sha256:6d306ce04f24ba4920ddc703124354c12b700c34f0cadd85be3e36db140ad143

Observation e022117d-be36-46b2-9812-c2b709538011 · outbound

This paper cites Almost optimal model-free reinforcement learning via reference-advantage decomposition.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Almost optimal model-free reinforcement learning via reference-advantage decomposition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.781590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.781590Z digest=sha256:e4d768658cb18e9ead91084018fc999908b689ed0493e245657c8ca5ba59d334

Observation 4bf24377-46c4-429f-a828-bbbcb4807ab2 · outbound

This paper cites The horizon-truncation lemma implies that the optimal H-step value is withinOpυqof the optimalH 1-step value.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The horizon-truncation lemma implies that the optimal H-step value is withinOpυqof the optimalH 1-step value

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.849097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.849097Z digest=sha256:b2beb69e0b36fb92f55263d03832d1a600002c763f0f566b34aa7b16d936e6a9

Observation 4f34ac99-42d8-4c7e-927e-d1aa34d12138 · outbound

This paper cites The proof is a backward induction on h.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The proof is a backward induction on h

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.932895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.932895Z digest=sha256:5969d30d239aa5f3149b6d2a6aeed0da99d31d5f0ed1a49b220bd1285bb9f00f

Observation 65ed91f3-fcdf-41d2-8e0d-e25b251464ab · outbound

This paper cites The process is stopped when the trajectory reaches the unlearned set Ok or when the frozen count of a learned pair doubles.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The process is stopped when the trajectory reaches the unlearned set Ok or when the frozen count of a learned pair doubles

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.996404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.996404Z digest=sha256:9a38a34b4491adfbecda2149725ffcd66eb9abca0a2b33e0dbb5bc90ae1f87cb

Observation 7d4e9a66-26ff-4174-b235-29c98665c644 · outbound

This paper cites This requires a variance closure argument for the optimistic values and the optimistic gaps.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence This requires a variance closure argument for the optimistic values and the optimistic gaps

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.111719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.111719Z digest=sha256:8a0115088ed8ad6eb6915e0cc92539da46718846af053b4f188a8e5030891fd1

Observation 56d6f6be-6c0f-41bb-ac7e-04de21c57a2d · outbound

This paper cites V ˚ H1`1.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence V ˚ H1`1

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.181643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.181643Z digest=sha256:76405de9746784da5c88473f8df4094ba24fce94ff5283523a54ba1fd822faba

Pith citing papers

No inbound Pith citation observations are available.