Pith. sign in

Paper Citation Record · LEDGER

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

As of 10 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2606.31769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31769 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy84
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14d4c429-7d32-4c9f-8a3e-03f50fc6269c · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.239431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:aca2c6fd134760a5ffadac69ea9058ee0396c6e7a5cb606cc7fe1995bd3182bd

Observation 73453eab-37e1-4f8b-967b-65daaab3520b · outbound

This paper cites Conference on Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Conference on Learning Theory , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.381207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cff2d5a871df8a63521047edef1d09ba9fe9c2e62fe8587e641431ded16df7ca

Observation 5db8998b-34fe-4760-8bcb-c3c0c5d4d7c8 · outbound

This paper cites Near-optimal Regret Using Policy Optimization in Online.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal Regret Using Policy Optimization in Online

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ec2a501f2515805f54498c9b1a82497339ee74028bfb7357dfbdf81597b3ba6e

Observation 5c621f20-e926-409d-ac62-4cb6a505d1c8 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.345951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:07288d342af3889f24419e853bc903deac01c077845d96e9cf5ada05ed5d1329

Observation 05a895cb-c2dc-4d63-b849-3a6d1a0eb3f3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.339102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e19a5655b4465414d2930a1f6bff976539c2f13c717322b47b321b0b12833a45

Observation 20323228-acb9-4341-9b07-f2fe5683d9ed · outbound

This paper cites Proceedings of the 34th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 34th International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.334006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:850d76765b8859d5e5c1d4c3964da6fb076f0b4c20c4e7907331fb3346987965

Observation b571c96a-376b-4eaf-bbdc-16a809693033 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.269768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9ac2c86b24bb498d7586c72037e52999720dd4b7fcc54f26a3655e90098b1b4c

Observation 85d2836e-8a92-4c80-b93e-fb39356e84a9 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.326980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:38c70e3d864b85934803ba42d26059b9d5cb92fb72b8fb4e77e0ffa0c426d626

Observation 3225afe0-49ef-4e5b-8a53-0a08dec03ee8 · outbound

This paper cites Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.397181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:91a67944b382e4975e929861d5a39eb07b5bbace938f3170edb41eaea49f0ee8

Observation 9ccd5974-909c-460a-b790-a20d1b0c7391 · outbound

This paper cites Non-Asymptotic Gap-Dependent Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Non-Asymptotic Gap-Dependent Regret Bounds for Tabular

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.361225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:267382a1d7b01e17d63854fac73db573041dcbc1c12267152f41a32797863995

Observation b960668c-5eb9-44c5-b74c-b90e2bfd2736 · outbound

This paper cites Online learning in episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online learning in episodic

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.359031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:513db90aa4b3d6d91b957bdb76c08320400764d61c815d9cf90e32c1120ef518

Observation 5a1fb6b0-2645-4226-a0fc-8d077f6a329c · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.319840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ed2a0fbe2fbdf1ce1913150f5101c983aa41b9ce7b739f7ab89b8ae7109514f1

Observation 45594917-e06d-48fe-81ec-bc38916cfaed · outbound

This paper cites Learning Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning Adversarial

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.268635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7cf280bb3890208608c44d69081990a75a9ab2e5e4554075e0aa9b1b06068c94

Observation 1757f675-696d-435d-adb6-25d277b97c66 · outbound

This paper cites Policy Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Policy Optimization in Adversarial

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.316384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:018e84fa898da8f3ae1aaf1594695fc4ded892e3ffc3bbf536523351a28097d4

Observation f194eb67-45c9-4105-b804-02d92d4f7fab · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.370981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d864745b0ea8a2a7dd97a55fdf7e72dec01b832056a3831fe2ac8486cb319d9f

Observation 6eb8e578-5ca1-4ca1-948b-67c146c569f9 · outbound

This paper cites Laurent and P.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Laurent and P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.312811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:26c4b5b15d710386d7e153d68bc3173768bc64404c556bee62ef93061dd9df1d

Observation c0992d9e-a455-4916-839e-cfc95685db02 · outbound

This paper cites Refined Lower Bounds for Adversarial Bandits , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Refined Lower Bounds for Adversarial Bandits , volume =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.259880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7017ce5e9fb8583df377965c64ea61a8ac6f43671160edd23db21279a139abef

Observation 0bdcf0ce-552d-4d47-8257-9bae3ad77771 · outbound

This paper cites Probability and Mathematical Statistics , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Probability and Mathematical Statistics , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.321775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:33be32cfa24c82a08358adcb8c28cf660fbfc36d258f4e76f70a3ec1e76be3b2

Observation 33c6b8d8-f7ab-4633-9ba0-64cd2dbc7cd0 · outbound

This paper cites 2013 , month =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2013 , month =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:619ed5e04e2929bf9443a9a115b3d3ca045aebaacf6baf376dd5763bcb09435a

Observation 1fa971e0-eb76-43ac-98ed-64bf069def4f · outbound

This paper cites Tuning Bandit Algorithms in Stochastic Environments.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tuning Bandit Algorithms in Stochastic Environments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.332323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:234d0db82f8e7184281cb91140d2b1abee4886d1ccb7f6790d5f9dfa8b3acb63

Observation dad72bcb-cd68-4dcc-a782-f0aab952156a · outbound

This paper cites Online Convex Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Convex Optimization in Adversarial

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a61451989b11fc920422fef3913a706ae314bf7e56194f0412b913e825da4e21

Observation eec37106-7870-4a44-9e85-da8a6caff019 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 37th International Conference on Machine Learning , pages =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4a4e0e78d1f9b51d35e55ee6d0221ce9a788033010375218ebea559adc3494cf

Observation d7911cad-e059-45e1-a737-dd470191f0d3 · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.223530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4bda6280e5bf478e4463d626ee8b0e5a9e9757fb9bc47f20e341b05e79e1533a

Observation 18fba09c-b30e-4ca9-9633-30eff050de10 · outbound

This paper cites Online Markov Decision Processes under Bandit Feedback , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Markov Decision Processes under Bandit Feedback , volume =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1c093c4b42aedd05dbca1e9f341dbecec72dcfc4fb9b6c46b7611cb7567f2974

Observation 3c4833ef-90c9-4662-99dc-7722f431363e · outbound

This paper cites Learning adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning adversarial

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.403917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:04477ec4a7423d0789869327a41cf4f78a1a481ee4e04479f61e39cfb0e20b59

Observation 8aeca06b-870a-4d02-9d44-e1de777d92ef · outbound

This paper cites Near-optimal regret for adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal regret for adversarial

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.297517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:af164f885980ab8efd020eec90479b8ac0a065cd38b2df98aad85e9f56d3e2db

Observation ef4a2f62-5098-4b17-9e9a-c3301233296b · outbound

This paper cites Delay-Adapted Policy Optimization and Improved Regret for Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Delay-Adapted Policy Optimization and Improved Regret for Adversarial

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.275251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:740f38ad4ca1e0bd13a1e2886c5a6829ebcf5a1484a95dff0c5eaad2d2963d9f

Observation 0e56bc8f-e625-4ed1-88c8-0fa5f1e6c559 · outbound

This paper cites Finite-Time Analysis of the Multiarmed Bandit Problem , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Finite-Time Analysis of the Multiarmed Bandit Problem , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.407915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:953d04d4fb9851a74a0e3bc220af060084c233c3554d158aeed8869148345d2a

Observation 55f4fc67-b198-4ae1-bcde-4649e38d84c1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.401904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c0c5a99745f7def7f90e394ba9a6442cde97a6fa2708a023cf9807a5124ebb55

Observation 79b11b62-41de-4c04-958e-5faa2952870a · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.301336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:93642f0a181437fa1b8ad9dec9e0ff9c89c97e0c41b0bcdb3caed7506e1690b4

Observation 61c2ad1d-8f56-4911-8464-6682442becd8 · outbound

This paper cites Simultaneously Learning Stochastic and Adversarial Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Simultaneously Learning Stochastic and Adversarial Episodic

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.314549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:5708804650903b53fca41680948dc10fdce931ca558e7509abdd35b2f4d5b5fd

Observation ead8046c-1c24-467a-bf8d-e9d780a30add · outbound

This paper cites The best of both worlds: stochastic and adversarial episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions The best of both worlds: stochastic and adversarial episodic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cecb0d4116882c7dcdb58a4e4237eddbd8934da5a22f714b98698e1f88bfcd85

Observation b7a99ceb-326d-47e8-8956-00025523569a · outbound

This paper cites Proceedings of the 31st Conference On Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st Conference On Learning Theory , pages =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.273202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0c30478c199b39ad54a5d1db89bc5641b4be5954123ca825364a50285ef6559d

Observation 7b6acfcb-4b3a-458a-82c3-a25c9792852a · outbound

This paper cites Tsallis-.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tsallis-

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1c145147e67001de4f66e3b59295a45d1b31c7a7af0a04566167ba1b7c9ed04b

Observation 283ef0d8-b7d3-4af1-88de-a7227d4cc57e · outbound

This paper cites Improved Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Analysis of the

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:68aebef60cd05f4bd96c7c5dd52416457a3aa536dba868d6f973720e5e397e53

Observation 1ad86420-6132-4b63-bdcf-211a82a68f68 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 36th International Conference on Machine Learning , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.335767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:89c65e429d7eb6179f04a0a33908a6eb62921f9c79ccbb5fb696e23736e41a31

Observation 069f8c57-4db0-4a62-9e12-88cdce021f82 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.385429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cb794cfed3f654b3a30bbf7feca9e5ef0bb325e607a6b6b5ed29162c97fd5df4

Observation f3a43589-aff7-4c0a-9851-de826d671e9d · outbound

This paper cites Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.309320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:890eb271f2e6f4cbf3a2cf1d753162ef71d61dfb68d6b1f97fe2ea5ec93a2a6d

Observation 65157022-e1de-4bf0-9fcc-b32a67f06990 · outbound

This paper cites Proceedings of The 27th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 27th Conference on Learning Theory , pages =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.289311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d76d83727a95b6269eec95eb89b1e7a8c0223320ac05610504b792c112eaf39a

Observation d29675b0-ca26-420e-8511-4cd3846ba2eb · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.230166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3595abd658451e7716163a9367a59576fdd549a66aa06f0fa581dc4821dd10f1

Observation b80e341a-a2d8-42a5-9890-bf1c7145cf5b · outbound

This paper cites 29th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 29th Annual Conference on Learning Theory , pages =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.295280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4122f497a90663ab0f68793d604ffd134e4cdf49c249c5a4161edc59b1af2b48

Observation 88bc35d7-fde4-467b-aa22-8d35e6dd8b77 · outbound

This paper cites An Improved Parametrization and Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions An Improved Parametrization and Analysis of the

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.237992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6df90677665bc2cf6bb2d4067fed1ee6c3c9324f8ed823ed0aeee76b4cd3080a

Observation 1fc76744-10bb-4bad-9298-0f5893c21a12 · outbound

This paper cites Adapting to Stochastic and Adversarial Losses in Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Adapting to Stochastic and Adversarial Losses in Episodic

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.266508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:17e21842e481bcc566e2e46ce6d0ae9c5bcf495197f27d3df8751dd98cc6bc6e

Observation 30f6b5b5-eda3-4ef6-a4b1-46800372c064 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.373563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e95f6429b5d34c822d822cf3638127a07f4a2df3ca71f55f4681e2ca5c706214

Observation ce6b0064-2ed7-49a0-b78b-04f67489f4c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.337481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:03d036b02e3a871d1ae1dc80d219f320481db98ebac87f09512e58fe445f0d7c

Observation ef46ec24-05b7-4742-ba9a-e616892f0806 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.402086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:06426e19153ebc458f4c03421b3bddf533c8e4a223efda8e7778205d211d54a3

Observation c8b63570-73c8-4d00-9521-052f06db5f20 · outbound

This paper cites Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:128f65e3f8ebc038885e5700bad1b90073e153d26f41c76584e18c1e69676b7c

Observation f761abec-1550-4e4b-8410-5c0bfe4d2265 · outbound

This paper cites Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.892402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:361f216770695ec425e864c9f74bcff6c852907076233570c46bae2223db7489

Observation b079470e-a09c-47a0-bd46-652dc5cec595 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Artificial Intelligence and Statistics , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.228386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:07b084d9fa6004e4146d414e38a50857c27bdfa2a24225f61e960b4c6cb2ab09

Observation 7351568b-25a9-4068-b13f-e3efda3a9fad · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:df2c0aaeee189b892fb68aa553c5bf2b9c1a297fbf1fe4d82769ce63dfaf97bf

Observation 975437ad-7c4b-4ea9-8ca0-5cc0b847ecbb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0fce61b318f7774570084e71a1eedaa96c162cb60d0ca20285eb233c6b7be137

Observation ebe4fd21-114e-47a6-9268-908d3f87c90c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.323736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2e6ca0f75007d00508d013361f592f224319c74c05c7c24bb8cb0390ce07c472

Observation 18ef48be-ef84-4f11-ac6e-5a0301de82d1 · outbound

This paper cites International Conference on Algorithmic Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Algorithmic Learning Theory , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f1686ee12032cc630a9c15bf2282fa0625f2eb27c9e36a6503c7f55d57a789b4

Observation f5336e48-97d2-452f-b844-47925bb868f2 · outbound

This paper cites Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.271271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f222e4241a3972bedd54baf1a24164d56902086053b1783c8223533d963b7a19

Observation 49eaa7ff-e90a-4c68-8843-76c717c97ab1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.248384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6db1256be3f679054e9fcbf748fca1594a989663f258592c3abc0b9439cba386

Observation 08eb070e-a267-4f22-a582-9e9212c0c8cb · outbound

This paper cites Proceedings of Algorithmic Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Algorithmic Learning Theory , pages =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.299423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f442a2a0d7ab4e09815780387e8651e8d73ce4520fcc4407a0480903c9f78e8e

Observation 0b9b304d-242e-44b7-9f9e-5a9ae06d0959 · outbound

This paper cites Proceedings of the Thirty-Second Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Thirty-Second Conference on Learning Theory , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.388785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ceb1c20524ab16e93456e2dbe29c4b28b0fcb69f98084b3777be2ff22a8b5532

Observation b2b371ee-f2dc-41f4-a3ae-5e9bcb46a8b2 · outbound

This paper cites Proceedings of the 28th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 28th Conference on Learning Theory , pages =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.390680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2d00c2483f772c131c599935970656f290be46168f7cd360a8d7d5f5e6de8ecc

Observation 2719bc76-2c44-4fdf-8959-5355e859fa57 · outbound

This paper cites Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.277184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e06d47c98a0a4b0df6ecff3aeb6562f43f1e73e45ed6775ff2de50753bec628c

Observation ad6ab9f6-3957-433f-bbf8-240cbc4db948 · outbound

This paper cites Gap-Dependent Bounds for.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Gap-Dependent Bounds for

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.330346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f508a86f79754b245f9ee6dd3efb89c3d19cb254cf9f94b27a91319747132557

Observation 734b30cd-ee00-4ded-b903-214fa1599dbd · outbound

This paper cites Proceedings of the 26th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 26th Annual Conference on Learning Theory , pages =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.258818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:5673c75303a25f491a9ce5dffb25e5326bcdcb58ea259ed3e3b2b6b9c64cf199

Observation 7a008eb7-8a0f-433c-aab6-24c3fe8865e9 · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.392809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:43b2efed79318408f589af76c436dedef37f29a064505037c213945be60b60dc

Observation 05a66b6a-d885-4d27-bf79-232c002494e2 · outbound

This paper cites SIAM Journal on Computing , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions SIAM Journal on Computing , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.264201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e812710abaf0ca8e624ef114a84e5c4dba110574a7009d7dd8015601b19655ac

Observation 16a038d2-a79f-469d-afdb-a9a93c832959 · outbound

This paper cites Proceedings of Thirty Eighth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Eighth Conference on Learning Theory , pages =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0aef0b6cc3fb506623ee5c14cfd4b6ea6e528c70b736887f502c699507df8a2e

Observation 1f334187-8f02-4e90-b1f0-4194950b9dd9 · outbound

This paper cites Proceedings of the Nineteenth International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Nineteenth International Conference on Machine Learning , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.291102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:59c31995bcd7dc1bf66af5439434e5082551990fb88b4ed4a56dc43e9026c7dd

Observation d1c1ef9b-7c8e-4ab8-87f0-75b1bede660e · outbound

This paper cites Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.386464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:dc259bc5e60c86f4f39d6ca11bf3835eef0296b091d1c4f694eeb2d4c582fbbb

Observation 46308e1a-e16c-4827-8310-11d5b19ee355 · outbound

This paper cites , author=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions , author=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:78410c3022292e26312b23bfecadb42a2128bea76b76f6f983c824a28c0a5b5c

Observation b9eb6bdc-d7d5-4fa2-9b68-cd3649abfeab · outbound

This paper cites IEEE Transactions on Neural Networks , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions IEEE Transactions on Neural Networks , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:93a33efb12791fe953690980838a87c5c92a755d085bb3f6a276c8a78c781e6e

Observation 56f0e67a-ed8d-45aa-aa73-4392dea58ee7 · outbound

This paper cites 2003 , booktitle =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2003 , booktitle =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.404206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ba56adb186cadbb057ae96a08c97abc0d4747a56d29cd76fb9cbf4f8af317fc6

Observation c2186f21-3fad-4ae9-a1c9-031e754cf68f · outbound

This paper cites Information and Computation , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Information and Computation , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7a138dfc11d69eab476f59ec51dd81cc17558ba3ac3bbde9ebfa2cc27889eaa1

Observation d0a6b86b-a741-4147-bdca-1198d3ff7a38 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proximal Policy Optimization Algorithms

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:55:35.889593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8019594b08bc95ff720283dfdb25bb2e73aa7903b6675d7e61b027610dc84864

Observation 5dac5bd2-ce00-4818-83e5-6edfbfc22023 · outbound

This paper cites and Veness, Joel and Bellemare, Marc G.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions and Veness, Joel and Bellemare, Marc G

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.256923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:938986bba05c4d51b666e592a3db3cc5f207316d222d18dbdc0f3857b55eb093

Observation 64ef28e4-d109-4630-8e76-61f152b6f25f · outbound

This paper cites Nature Medicine , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Nature Medicine , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.372950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8cdae61724dc514d48284550f42f43e7733a2af045680dc0e2efb643dc06b39f

Observation 6b2b7331-052d-422e-913b-d221591b1cf7 · outbound

This paper cites Empirical.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Empirical

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.328708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e1b0c6d0e2a113f90ac2f1ae3856c83a01ad4446d21bad6bc8d9b9e77ed5b124

Observation 02221255-3002-405d-9dea-7106dafe63e7 · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:87bad39964eec4062a622530802862bdffddbde77339929f3f900cb60476bc57

Observation 0e3c5536-7364-4b30-bac2-00b4f2c0ab60 · outbound

This paper cites Reinforcement Learning from Adversarial Preferences in Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Reinforcement Learning from Adversarial Preferences in Tabular

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.279149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2563c45f11b1953dee8c6b4b43b58fe3deecacb4e6098447b6e94e77e863214b

Observation 048475ad-6b6f-4e6d-82b6-f8446350cf73 · outbound

This paper cites Proceedings of Thirty Seventh Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Seventh Conference on Learning Theory , pages =

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.288156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0654bbcfc316322db4c933b2f3c28ade554de9bf390a14ef4e08d284a1a33d2d

Observation 328d5ee9-e3a3-488f-a8d2-9aea70f402ed · outbound

This paper cites M arkov Decision Processes with Arbitrary Reward Processes.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions M arkov Decision Processes with Arbitrary Reward Processes

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.307301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b4e28031e0ca62aca86d4f04f50c0444676bb10f98e9e926187a5a41399b2ea0

Observation 07fd084d-5d60-4008-855a-81960e3a7ce2 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.311024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:58a252813238a0b18036577f3be6f638826c76ab146dffa21762b45ae6b25b37

Observation a85d3468-c1fd-4e7b-b48d-8c1283dbe9b1 · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 39th International Conference on Machine Learning , pages =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.234774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:09502b477e58714953ed9174122ca010209f4f8bfe88b1906211989b7a976ffe

Observation e9096485-519a-4d85-8e7e-de228c345a8e · outbound

This paper cites 2009 , author =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2009 , author =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b64bcf39afd7d54a7b21206b229c5a5f6fffbcb40ffe21f15d17a9d2ad49b89d

Observation b800689a-e08e-43bc-b1d7-7f7352943bba · outbound

This paper cites Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.363486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f3aed0ffe1e5b88c92b9538e963a6385d618365007c198766dc680e344d18ce5

Observation abb9c9b2-1291-4586-9c3f-a45891fbcc19 · outbound

This paper cites Episodic Reinforcement Learning in Finite.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Episodic Reinforcement Learning in Finite

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.355915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3d4a407da47c7b7d097cc5cb53cb090e3c950cb4232820e170df9fe9dd237621

Observation 2604cc15-1e64-45f0-b27d-a2c119553618 · outbound

This paper cites Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d4f05d911b258a61505d6ad7dd98cf6aa58a442c3779bfadca10fba5544a4ded

Observation 6865871f-6166-4178-ad00-ffc61fb2db14 · outbound

This paper cites International Conference on Machine Learning , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , year =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.354755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:41005842c0f4ccc7f6a056d246640e1761ed1648a29cd6c1f160bd9d24ad90e3

Observation 5d6dfc3b-f803-42a3-a658-c37bc2f2bb6f · outbound

This paper cites Narrowing the Gap between Adversarial and Stochastic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Narrowing the Gap between Adversarial and Stochastic

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ae4eb32585f1bf4a05451e5173fb196ec0559a56deb771bcbe9efe53e1992219

Observation 40ba6c59-3631-4677-ba57-2d8ab9a7cc61 · outbound

This paper cites Proceedings of the 32nd International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 32nd International Conference on Machine Learning , pages =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.318135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8f52ba00a38979f62997a03614d5c0cc1a53e6aa7ebfa298410743a68c2e2ff3

Observation 1283848a-a0b0-449f-91e3-da4a95ad938c · outbound

This paper cites Advances in Neural Information Processing Systems , publisher =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , publisher =

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.343968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:fb2fab4c2dea5260113993e580fdc3dcb5f98b86895086377bd708d3d11cdbcf

Observation 65a0f318-5b1e-4248-81bb-e8d6426791b9 · outbound

This paper cites Fine-Grained Gap-Dependent Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Fine-Grained Gap-Dependent Bounds for Tabular

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:38805493daa914d00c58095d82be56e32a455753d32f28c5bb9866f9c8bbb3fa

Pith citing papers

No inbound Pith citation observations are available.