Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2606.31769.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 14d4c429-7d32-4c9f-8a3e-03f50fc6269c · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73453eab-37e1-4f8b-967b-65daaab3520b · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Conference on Learning Theory , pages=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5db8998b-34fe-4760-8bcb-c3c0c5d4d7c8 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal Regret Using Policy Optimization in Online
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c621f20-e926-409d-ac62-4cb6a505d1c8 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 41st International Conference on Machine Learning , pages=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05a895cb-c2dc-4d63-b849-3a6d1a0eb3f3 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20323228-acb9-4341-9b07-f2fe5683d9ed · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 34th International Conference on Machine Learning , pages =
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b571c96a-376b-4eaf-bbdc-16a809693033 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85d2836e-8a92-4c80-b93e-fb39356e84a9 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3225afe0-49ef-4e5b-8a53-0a08dec03ee8 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ccd5974-909c-460a-b790-a20d1b0c7391 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Non-Asymptotic Gap-Dependent Regret Bounds for Tabular
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b960668c-5eb9-44c5-b74c-b90e2bfd2736 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online learning in episodic
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a1fb6b0-2645-4226-a0fc-8d077f6a329c · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45594917-e06d-48fe-81ec-bc38916cfaed · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning Adversarial
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1757f675-696d-435d-adb6-25d277b97c66 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Policy Optimization in Adversarial
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f194eb67-45c9-4105-b804-02d92d4f7fab · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6eb8e578-5ca1-4ca1-948b-67c146c569f9 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Laurent and P
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0992d9e-a455-4916-839e-cfc95685db02 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Refined Lower Bounds for Adversarial Bandits , volume =
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bdcf0ce-552d-4d47-8257-9bae3ad77771 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Probability and Mathematical Statistics , volume=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 33c6b8d8-f7ab-4633-9ba0-64cd2dbc7cd0 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2013 , month =
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fa971e0-eb76-43ac-98ed-64bf069def4f · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tuning Bandit Algorithms in Stochastic Environments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dad72bcb-cd68-4dcc-a782-f0aab952156a · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Convex Optimization in Adversarial
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eec37106-7870-4a44-9e85-da8a6caff019 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 37th International Conference on Machine Learning , pages =
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d7911cad-e059-45e1-a737-dd470191f0d3 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18fba09c-b30e-4ca9-9633-30eff050de10 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Markov Decision Processes under Bandit Feedback , volume =
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3c4833ef-90c9-4662-99dc-7722f431363e · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning adversarial
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8aeca06b-870a-4d02-9d44-e1de777d92ef · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal regret for adversarial
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ef4a2f62-5098-4b17-9e9a-c3301233296b · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Delay-Adapted Policy Optimization and Improved Regret for Adversarial
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e56bc8f-e625-4ed1-88c8-0fa5f1e6c559 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Finite-Time Analysis of the Multiarmed Bandit Problem , year =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55f4fc67-b198-4ae1-bcde-4649e38d84c1 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79b11b62-41de-4c04-958e-5faa2952870a · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61c2ad1d-8f56-4911-8464-6682442becd8 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Simultaneously Learning Stochastic and Adversarial Episodic
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ead8046c-1c24-467a-bf8d-e9d780a30add · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions The best of both worlds: stochastic and adversarial episodic
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7a99ceb-326d-47e8-8956-00025523569a · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st Conference On Learning Theory , pages =
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b6acfcb-4b3a-458a-82c3-a25c9792852a · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tsallis-
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 283ef0d8-b7d3-4af1-88de-a7227d4cc57e · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Analysis of the
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ad86420-6132-4b63-bdcf-211a82a68f68 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 36th International Conference on Machine Learning , pages =
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 069f8c57-4db0-4a62-9e12-88cdce021f82 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3a43589-aff7-4c0a-9851-de826d671e9d · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 65157022-e1de-4bf0-9fcc-b32a67f06990 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 27th Conference on Learning Theory , pages =
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d29675b0-ca26-420e-8511-4cd3846ba2eb · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b80e341a-a2d8-42a5-9890-bf1c7145cf5b · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 29th Annual Conference on Learning Theory , pages =
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 88bc35d7-fde4-467b-aa22-8d35e6dd8b77 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions An Improved Parametrization and Analysis of the
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fc76744-10bb-4bad-9298-0f5893c21a12 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Adapting to Stochastic and Adversarial Losses in Episodic
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30f6b5b5-eda3-4ef6-a4b1-46800372c064 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce6b0064-2ed7-49a0-b78b-04f67489f4c6 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ef46ec24-05b7-4742-ba9a-e616892f0806 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8b63570-73c8-4d00-9521-052f06db5f20 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f761abec-1550-4e4b-8410-5c0bfe4d2265 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b079470e-a09c-47a0-bd46-652dc5cec595 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Artificial Intelligence and Statistics , pages=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7351568b-25a9-4068-b13f-e3efda3a9fad · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 38th International Conference on Machine Learning , pages =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 975437ad-7c4b-4ea9-8ca0-5cc0b847ecbb · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ebe4fd21-114e-47a6-9268-908d3f87c90c · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18ef48be-ef84-4f11-ac6e-5a0301de82d1 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Algorithmic Learning Theory , pages=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5336e48-97d2-452f-b844-47925bb868f2 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49eaa7ff-e90a-4c68-8843-76c717c97ab1 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08eb070e-a267-4f22-a582-9e9212c0c8cb · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Algorithmic Learning Theory , pages =
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b9b304d-242e-44b7-9f9e-5a9ae06d0959 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Thirty-Second Conference on Learning Theory , pages =
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2b371ee-f2dc-41f4-a3ae-5e9bcb46a8b2 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 28th Conference on Learning Theory , pages =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2719bc76-2c44-4fdf-8959-5355e859fa57 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad6ab9f6-3957-433f-bbf8-240cbc4db948 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Gap-Dependent Bounds for
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 734b30cd-ee00-4ded-b903-214fa1599dbd · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 26th Annual Conference on Learning Theory , pages =
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7a008eb7-8a0f-433c-aab6-24c3fe8865e9 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05a66b6a-d885-4d27-bf79-232c002494e2 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions SIAM Journal on Computing , year =
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16a038d2-a79f-469d-afdb-a9a93c832959 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Eighth Conference on Learning Theory , pages =
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f334187-8f02-4e90-b1f0-4194950b9dd9 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Nineteenth International Conference on Machine Learning , pages=
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d1c1ef9b-7c8e-4ab8-87f0-75b1bede660e · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 46308e1a-e16c-4827-8310-11d5b19ee355 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions , author=
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9eb6bdc-d7d5-4fa2-9b68-cd3649abfeab · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions IEEE Transactions on Neural Networks , volume=
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 56f0e67a-ed8d-45aa-aa73-4392dea58ee7 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2003 , booktitle =
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2186f21-3fad-4ae9-a1c9-031e754cf68f · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Information and Computation , volume=
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0a6b86b-a741-4147-bdca-1198d3ff7a38 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proximal Policy Optimization Algorithms
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5dac5bd2-ce00-4818-83e5-6edfbfc22023 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions and Veness, Joel and Bellemare, Marc G
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64ef28e4-d109-4630-8e76-61f152b6f25f · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Nature Medicine , volume=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6b2b7331-052d-422e-913b-d221591b1cf7 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Empirical
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02221255-3002-405d-9dea-7106dafe63e7 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e3c5536-7364-4b30-bac2-00b4f2c0ab60 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Reinforcement Learning from Adversarial Preferences in Tabular
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 048475ad-6b6f-4e6d-82b6-f8446350cf73 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Seventh Conference on Learning Theory , pages =
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 328d5ee9-e3a3-488f-a8d2-9aea70f402ed · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions M arkov Decision Processes with Arbitrary Reward Processes
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07fd084d-5d60-4008-855a-81960e3a7ce2 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a85d3468-c1fd-4e7b-b48d-8c1283dbe9b1 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 39th International Conference on Machine Learning , pages =
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9096485-519a-4d85-8e7e-de228c345a8e · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2009 , author =
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b800689a-e08e-43bc-b1d7-7f7352943bba · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abb9c9b2-1291-4586-9c3f-a45891fbcc19 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Episodic Reinforcement Learning in Finite
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2604cc15-1e64-45f0-b27d-a2c119553618 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6865871f-6166-4178-ad00-ffc61fb2db14 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , year =
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d6dfc3b-f803-42a3-a658-c37bc2f2bb6f · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Narrowing the Gap between Adversarial and Stochastic
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40ba6c59-3631-4677-ba57-2d8ab9a7cc61 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 32nd International Conference on Machine Learning , pages =
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1283848a-a0b0-449f-91e3-da4a95ad938c · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , publisher =
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 65a0f318-5b1e-4248-81bb-e8d6426791b9 · outbound
Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Fine-Grained Gap-Dependent Bounds for Tabular
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.