Pith. sign in

Paper Citation Record · LEDGER

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions

As of 10 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2606.31769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31769 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:55:27.313461Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy84
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14d4c429-7d32-4c9f-8a3e-03f50fc6269c · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.239431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a1e2dd1161d45ce32b196b4f937d4ff32f163437ef640a312dc7c29f7aaf8099

Observation 73453eab-37e1-4f8b-967b-65daaab3520b · outbound

This paper cites Conference on Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Conference on Learning Theory , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.381207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:813475be10159fae4d5b8424a0979eaa1a06976bfeb4bdb6910ba9834cfe8647

Observation 5db8998b-34fe-4760-8bcb-c3c0c5d4d7c8 · outbound

This paper cites Near-optimal Regret Using Policy Optimization in Online.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal Regret Using Policy Optimization in Online

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.341134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e832624b2c31e2be5515098d7a8f2294d7e381a229def3fa8200b3edef10cf4b

Observation 5c621f20-e926-409d-ac62-4cb6a505d1c8 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.345951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:4577a782d54da8b5268a4adea37ecb4946881a7ff6be5fc0aab4ea73a2f44779

Observation 05a895cb-c2dc-4d63-b849-3a6d1a0eb3f3 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.339102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:8d9d5886ed9afda06a2c659f5283553b48edc8a2f7dffc4aceb29c7d2e666b46

Observation 20323228-acb9-4341-9b07-f2fe5683d9ed · outbound

This paper cites Proceedings of the 34th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 34th International Conference on Machine Learning , pages =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.334006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:52e6ca8ebbcfe57f9c8f4cb19ddfc09e8cb80a304ea8a830ea1b5af7c062fa08

Observation b571c96a-376b-4eaf-bbdc-16a809693033 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.269768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d8f72bf9e526e2076837db374783c6f823bf4931aadf16e07937d8918ce5893c

Observation 85d2836e-8a92-4c80-b93e-fb39356e84a9 · outbound

This paper cites International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , pages=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.326980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:606ea2faaaaf456faf89d233ce2a8617441998c6ded639a7a4389ade22ecdc6d

Observation 3225afe0-49ef-4e5b-8a53-0a08dec03ee8 · outbound

This paper cites Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.397181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:c0df4e284b3173a1017d9fe301d228d2aa6de2b859eb04e02ffee3d5e11626c7

Observation 9ccd5974-909c-460a-b790-a20d1b0c7391 · outbound

This paper cites Non-Asymptotic Gap-Dependent Regret Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Non-Asymptotic Gap-Dependent Regret Bounds for Tabular

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.361225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7311c988bd1a0e1a2c02132419284a6d73814c5fbf5c409a013263c56428d495

Observation b960668c-5eb9-44c5-b74c-b90e2bfd2736 · outbound

This paper cites Online learning in episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online learning in episodic

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.359031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:272275aabeff2b63acdb6595096a820e24f967587edd6e90f8a27d3521a8bc7a

Observation 5a1fb6b0-2645-4226-a0fc-8d077f6a329c · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.319840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:006493f44a888fa5f467e8892b3487b23f0a6d6dc7893ebf87f478bf9c1ec33c

Observation 45594917-e06d-48fe-81ec-bc38916cfaed · outbound

This paper cites Learning Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning Adversarial

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.268635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b8644637e308d600276c8e57880ee804f2211587795a053f2dd1a8ed85ffe702

Observation 1757f675-696d-435d-adb6-25d277b97c66 · outbound

This paper cites Policy Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Policy Optimization in Adversarial

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.316384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9bea6de96ce2bd1e3364fe65867366a15d7737f30c8181b776a6ccac18fcfc08

Observation f194eb67-45c9-4105-b804-02d92d4f7fab · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.370981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1a7339800b217856c16758d43a5d96995c3461103ab113f23e9348f9293a64cd

Observation 6eb8e578-5ca1-4ca1-948b-67c146c569f9 · outbound

This paper cites Laurent and P.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Laurent and P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.312811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2ef5eb2840ef0bbe0f3a8e8e9f6c0e687f2fe77883974a55b525aef579ec00bc

Observation c0992d9e-a455-4916-839e-cfc95685db02 · outbound

This paper cites Refined Lower Bounds for Adversarial Bandits , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Refined Lower Bounds for Adversarial Bandits , volume =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.259880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ad5f09e8caf9d853887f0cc004fb1bde462248fdaebc9998992e291171bab86a

Observation 0bdcf0ce-552d-4d47-8257-9bae3ad77771 · outbound

This paper cites Probability and Mathematical Statistics , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Probability and Mathematical Statistics , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.321775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:657d1ab5082b49d5785f95edcbf409e3de92d4ff8aff12bd096ac8d57a3f12e2

Observation 33c6b8d8-f7ab-4633-9ba0-64cd2dbc7cd0 · outbound

This paper cites 2013 , month =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2013 , month =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:829cf2ac5a4503b16fa283a0f874eb5381fefaa310fbd6a2079ce927131c85ff

Observation 1fa971e0-eb76-43ac-98ed-64bf069def4f · outbound

This paper cites Tuning Bandit Algorithms in Stochastic Environments.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tuning Bandit Algorithms in Stochastic Environments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.332323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b30a739b596c94c2397442bc36d968d74a7cdec8ce6c2f5bc9018bf4b3b6ebc3

Observation dad72bcb-cd68-4dcc-a782-f0aab952156a · outbound

This paper cites Online Convex Optimization in Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Convex Optimization in Adversarial

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1dab25afc94242689ccf139091a49000bc43f8d2c77d963d4ead4939e71fb88a

Observation eec37106-7870-4a44-9e85-da8a6caff019 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 37th International Conference on Machine Learning , pages =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:95a3558b998e193cc8b929017c78d081d840e69632fcab3cbb27c617c1be9b58

Observation d7911cad-e059-45e1-a737-dd470191f0d3 · outbound

This paper cites an unresolved cited work.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:52:40.223530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ad344dbff8a0c899fdc70cdc192599ca9a056057119a1e6ab0ebe968941201cf

Observation 18fba09c-b30e-4ca9-9633-30eff050de10 · outbound

This paper cites Online Markov Decision Processes under Bandit Feedback , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Online Markov Decision Processes under Bandit Feedback , volume =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a65f59731a564735a46403377d5e77f33d80bac016d1eb2ffc7bd18b26f536a4

Observation 3c4833ef-90c9-4662-99dc-7722f431363e · outbound

This paper cites Learning adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Learning adversarial

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.403917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:059e83193baa4cdf2e33fced26abe3a15f591c7a7ab56daf0326c24ec414cbf9

Observation 8aeca06b-870a-4d02-9d44-e1de777d92ef · outbound

This paper cites Near-optimal regret for adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Near-optimal regret for adversarial

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.297517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:f690c87e91748c466ecbf979c94b78f6d91e5a2f5d9c9a163e850cb090de8fa9

Observation ef4a2f62-5098-4b17-9e9a-c3301233296b · outbound

This paper cites Delay-Adapted Policy Optimization and Improved Regret for Adversarial.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Delay-Adapted Policy Optimization and Improved Regret for Adversarial

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.275251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9f5b4ea75f7690e4e4a91c2c31880766b1b29a4fb81a915cd0bcfe1062aea9bf

Observation 0e56bc8f-e625-4ed1-88c8-0fa5f1e6c559 · outbound

This paper cites Finite-Time Analysis of the Multiarmed Bandit Problem , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Finite-Time Analysis of the Multiarmed Bandit Problem , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.407915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:215b99a802a4ac5660c8fe0ca048603647ad9d70784258cc81688bf3d6aad952

Observation 55f4fc67-b198-4ae1-bcde-4649e38d84c1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.401904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9341c2fa4e0c2d1dfeed0798f438ec55dbdf1e23b4a7b3eb44232068aa94e1db

Observation 79b11b62-41de-4c04-958e-5faa2952870a · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.301336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9fe7695972492a551853f8f790fa56a646181e309e75f78a837709f17cf633ea

Observation 61c2ad1d-8f56-4911-8464-6682442becd8 · outbound

This paper cites Simultaneously Learning Stochastic and Adversarial Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Simultaneously Learning Stochastic and Adversarial Episodic

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.314549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1c610b094b1d391c4247e50213a1f68aed05757a15c873f3799634de6c702762

Observation ead8046c-1c24-467a-bf8d-e9d780a30add · outbound

This paper cites The best of both worlds: stochastic and adversarial episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions The best of both worlds: stochastic and adversarial episodic

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e61b117da01aae6101faccca065c66bf6e2421f23c928cc17b05e9adc7e2b63d

Observation b7a99ceb-326d-47e8-8956-00025523569a · outbound

This paper cites Proceedings of the 31st Conference On Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st Conference On Learning Theory , pages =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.273202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:bab9b928dfef95857c1c7e058d9f15a6cf9d50aa333e647a6fdd1efee08385cc

Observation 7b6acfcb-4b3a-458a-82c3-a25c9792852a · outbound

This paper cites Tsallis-.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Tsallis-

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.395203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0e4ae155bdf0b764e72d7fd6b1aa1167e01562621cebd98847dc8441c63dd89d

Observation 283ef0d8-b7d3-4af1-88de-a7227d4cc57e · outbound

This paper cites Improved Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Analysis of the

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0710135fc08a1bb6af3cf81938a18b26713800ce47c34e1be2223eeea132c3cb

Observation 1ad86420-6132-4b63-bdcf-211a82a68f68 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 36th International Conference on Machine Learning , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.335767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:389db52c0ceabeb33b3c98cfab32ede77114362792c511718324ec2eb7287e03

Observation 069f8c57-4db0-4a62-9e12-88cdce021f82 · outbound

This paper cites Proceedings of Thirty Fourth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Fourth Conference on Learning Theory , pages =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.385429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:40af6ff874f053a6b5db3314805913cf17cbd5f3cf5ac27deb892172726158b4

Observation f3a43589-aff7-4c0a-9851-de826d671e9d · outbound

This paper cites Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Towards Best-of-All-Worlds Online Learning with Feedback Graphs , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.309320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:cf58be45f2f0b45b19868f058de6ae16fb14125535673146d6b0db5fca83974a

Observation 65157022-e1de-4bf0-9fcc-b32a67f06990 · outbound

This paper cites Proceedings of The 27th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 27th Conference on Learning Theory , pages =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.289311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:731163261beac394bb40e753693433d3df1df787cd3bbce96c58dffdfdb2ccda

Observation d29675b0-ca26-420e-8511-4cd3846ba2eb · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.230166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:87182ade8c0ca10992ddc01dd9453d5aabcb67222c43255ed5737e76f5d8f707

Observation b80e341a-a2d8-42a5-9890-bf1c7145cf5b · outbound

This paper cites 29th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 29th Annual Conference on Learning Theory , pages =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.295280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:41dd237a03061483c582ddbaa994787c0703d87aa8de10df0c617f59bbc85312

Observation 88bc35d7-fde4-467b-aa22-8d35e6dd8b77 · outbound

This paper cites An Improved Parametrization and Analysis of the.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions An Improved Parametrization and Analysis of the

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.237992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1e9be0821a8166a4f03f2c721461ac3eade423da98dca3cea9fd8aadcd193710

Observation 1fc76744-10bb-4bad-9298-0f5893c21a12 · outbound

This paper cites Adapting to Stochastic and Adversarial Losses in Episodic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Adapting to Stochastic and Adversarial Losses in Episodic

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.266508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:99fb0b5220dc955efa5f6adb234dabf2b0325df4d14f5a02589c047e244c4693

Observation 30f6b5b5-eda3-4ef6-a4b1-46800372c064 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.373563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:91a2b6f6984f129707d333d946bcf25096bfddbf47721a8e44d3b697dcac7aa6

Observation ce6b0064-2ed7-49a0-b78b-04f67489f4c6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.337481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0e4a0355e974d834d8fe4cf3d215920e1f476651635bb1d2bb059504227fdef0

Observation ef46ec24-05b7-4742-ba9a-e616892f0806 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.402086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a0a90589ca8a288fb5c10dc9310ac2da86a360f2d2a4b810e9deaa8054cd71b3

Observation c8b63570-73c8-4d00-9521-052f06db5f20 · outbound

This paper cites Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:02642f8c11b581dc60397c2a82bc92ec244bb102b15608944d70fe324b4c8893

Observation f761abec-1550-4e4b-8410-5c0bfe4d2265 · outbound

This paper cites Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.892402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:2c6fd38724e1864734457f9c049d24a81774dab5b2436e4e3c24eceff0f36714

Observation b079470e-a09c-47a0-bd46-652dc5cec595 · outbound

This paper cites International Conference on Artificial Intelligence and Statistics , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Artificial Intelligence and Statistics , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.228386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b63f908a8c520f781fea78e9db80741679b4d6bf3ea793f572fb6a5aa079064a

Observation 7351568b-25a9-4068-b13f-e3efda3a9fad · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.406228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9dd0ba355a1954863ecc082946c069cb633342ae6d0fd859a627f62796fa28d5

Observation 975437ad-7c4b-4ea9-8ca0-5cc0b847ecbb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e077f2b71cee071048ac6eeae33f1ced74d4d48f0b362611dc324af7f19d110d

Observation ebe4fd21-114e-47a6-9268-908d3f87c90c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.323736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:dc19571b5c7b7f86d4e54b6c9e759aac167237ec5ff8e468decaa0dd45243fbf

Observation 18ef48be-ef84-4f11-ac6e-5a0301de82d1 · outbound

This paper cites International Conference on Algorithmic Learning Theory , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Algorithmic Learning Theory , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1d76178821124333f8cecabdaaaf4696af5af5f266ee1331b74302bca117ef94

Observation f5336e48-97d2-452f-b844-47925bb868f2 · outbound

This paper cites Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Bias no more: high-probability data-dependent regret bounds for adversarial bandits and MDPs , volume =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.271271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b7f60ba68ab12520c3dcb523a9a4edaba39f571551489e12c392d423679ce7b9

Observation 49eaa7ff-e90a-4c68-8843-76c717c97ab1 · outbound

This paper cites Journal of Machine Learning Research , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Journal of Machine Learning Research , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.248384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:00d7ab60069704f2d81c465344f5360e122750c4033a801e9587858313a0981b

Observation 08eb070e-a267-4f22-a582-9e9212c0c8cb · outbound

This paper cites Proceedings of Algorithmic Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Algorithmic Learning Theory , pages =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.299423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7683fb0754a40e368babf1b1a2ff9be33c2a1c73514ae62d05ccfcc8a717a884

Observation 0b9b304d-242e-44b7-9f9e-5a9ae06d0959 · outbound

This paper cites Proceedings of the Thirty-Second Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Thirty-Second Conference on Learning Theory , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.388785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e00d2e6d852d067b926c32b00dcd83778af31e2ac5207ca7e70dfdc38ed5a95d

Observation b2b371ee-f2dc-41f4-a3ae-5e9bcb46a8b2 · outbound

This paper cites Proceedings of the 28th Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 28th Conference on Learning Theory , pages =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.390680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6f02450ab28bac22e1d4013efa1bb87ed4a6f6dcb08a1af78cf9e45c6a83f45a

Observation 2719bc76-2c44-4fdf-8959-5355e859fa57 · outbound

This paper cites Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits:

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.277184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:994d7927dc06c121391f084e426515b880dee74fe6e00d019f6ea34bffa4c92f

Observation ad6ab9f6-3957-433f-bbf8-240cbc4db948 · outbound

This paper cites Gap-Dependent Bounds for.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Gap-Dependent Bounds for

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.330346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:3dec15c25d8bc770cb8b1fde50f68feddb69ae0bee30313ce87a08ddf3f3f0e9

Observation 734b30cd-ee00-4ded-b903-214fa1599dbd · outbound

This paper cites Proceedings of the 26th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 26th Annual Conference on Learning Theory , pages =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.258818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1705dbc47eaea82850948776e01d2abf7602e21b65dbe70e6dc3e301c789b13f

Observation 7a008eb7-8a0f-433c-aab6-24c3fe8865e9 · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 31st International Conference on Machine Learning , pages =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.392809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:020dcbf7ec89438f53443f4427fd8a38d1a8f824b45a69fb52369c8fa9279078

Observation 05a66b6a-d885-4d27-bf79-232c002494e2 · outbound

This paper cites SIAM Journal on Computing , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions SIAM Journal on Computing , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.264201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7398029039d8d4eb5cbaae571336cd08d554f175f6d2e9f9bdb566a9c705eea9

Observation 16a038d2-a79f-469d-afdb-a9a93c832959 · outbound

This paper cites Proceedings of Thirty Eighth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Eighth Conference on Learning Theory , pages =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e0d6e4f28d5360454efbcccc0302f2af6c006a41dc73bab5c7395957b48370de

Observation 1f334187-8f02-4e90-b1f0-4194950b9dd9 · outbound

This paper cites Proceedings of the Nineteenth International Conference on Machine Learning , pages=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Nineteenth International Conference on Machine Learning , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.291102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:71ba6d88676c8a7a934b3222bcd0f6935928bb9a35e1e0b392a942996812ad04

Observation d1c1ef9b-7c8e-4ab8-87f0-75b1bede660e · outbound

This paper cites Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fifteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.386464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:78c592b576f016b33c22177a378d4efa6fa510579487137a02435e7be2be09ab

Observation 46308e1a-e16c-4827-8310-11d5b19ee355 · outbound

This paper cites , author=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions , author=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.378348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:215214276449101265502cbc07da1eca9d5311b8d1c5ce3a5e236e044bf5f371

Observation b9eb6bdc-d7d5-4fa2-9b68-cd3649abfeab · outbound

This paper cites IEEE Transactions on Neural Networks , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions IEEE Transactions on Neural Networks , volume=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.325515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b9f2ccebfcbf259356ac54168a37a3f462682d0737661ae35829dc44d080f1cb

Observation 56f0e67a-ed8d-45aa-aa73-4392dea58ee7 · outbound

This paper cites 2003 , booktitle =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2003 , booktitle =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.404206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:062228138cab14d90d65cfe24d4ebe69fc53fba4f4645e54de7b496b211eb8a7

Observation c2186f21-3fad-4ae9-a1c9-031e754cf68f · outbound

This paper cites Information and Computation , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Information and Computation , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.352052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d7986105ed290257dfb8f955ee7e5ab9bb61213be231a60d0711b9820e01b284

Observation d0a6b86b-a741-4147-bdca-1198d3ff7a38 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proximal Policy Optimization Algorithms

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:55:35.889593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:36102d1bee99ddd8f81e1f76e70fd85ca469765c97f47dea6072dc09dac11964

Observation 5dac5bd2-ce00-4818-83e5-6edfbfc22023 · outbound

This paper cites and Veness, Joel and Bellemare, Marc G.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions and Veness, Joel and Bellemare, Marc G

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.256923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:e2ee8f10679790d4160219d15fbb3f4c0c47ffd70e102c7f7a9ade3d5d9bcc9b

Observation 64ef28e4-d109-4630-8e76-61f152b6f25f · outbound

This paper cites Nature Medicine , volume=.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Nature Medicine , volume=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.372950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:58105a58c2b8f2fad6e0fffb6064533180d14042f9ecd5f91caf9fad7eb15026

Observation 6b2b7331-052d-422e-913b-d221591b1cf7 · outbound

This paper cites Empirical.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Empirical

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.328708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:110b6d008f5295476e012a01dc988e9ba987a6fa009f9468fb9b3d8edbaebb82

Observation 02221255-3002-405d-9dea-7106dafe63e7 · outbound

This paper cites Proceedings of the 25th Annual Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 25th Annual Conference on Learning Theory , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d8ff1f99cac39830a9586598645bb26a33380f3ea8d1b6063d4b371b5045187c

Observation 0e3c5536-7364-4b30-bac2-00b4f2c0ab60 · outbound

This paper cites Reinforcement Learning from Adversarial Preferences in Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Reinforcement Learning from Adversarial Preferences in Tabular

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.279149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:7d3eca0ec5146eb82e65c47b47adbb909453bc069a7b1ea17e1bcc803986a1c4

Observation 048475ad-6b6f-4e6d-82b6-f8446350cf73 · outbound

This paper cites Proceedings of Thirty Seventh Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Seventh Conference on Learning Theory , pages =

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.288156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:9559c71ff3cf9705c4f1e2a314e7ccf65afc8601c43b776460b0cb96cf054e1a

Observation 328d5ee9-e3a3-488f-a8d2-9aea70f402ed · outbound

This paper cites M arkov Decision Processes with Arbitrary Reward Processes.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions M arkov Decision Processes with Arbitrary Reward Processes

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.307301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:ed7b9c5d651c6d46f9d6aae033ada3ffc0866f7202988e8edd2f49afa75323a6

Observation 07fd084d-5d60-4008-855a-81960e3a7ce2 · outbound

This paper cites Proceedings of Thirty Sixth Conference on Learning Theory , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of Thirty Sixth Conference on Learning Theory , pages =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.311024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:891a76a2536e0c370c3076f8f222dfd1c623f4fe4804c704f9cca8dde00fc83d

Observation a85d3468-c1fd-4e7b-b48d-8c1283dbe9b1 · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 39th International Conference on Machine Learning , pages =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.234774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:a04bbc5352954907b2a467e777594b114fc4bd3922b5815c4f93088615c83597

Observation e9096485-519a-4d85-8e7e-de228c345a8e · outbound

This paper cites 2009 , author =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions 2009 , author =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.305606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:1f9bf36aa487c8b92cc26a11b923858b9ae6a999e7de8c90c92d90222de922aa

Observation b800689a-e08e-43bc-b1d7-7f7352943bba · outbound

This paper cites Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds , volume =

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.363486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:26c2cf491e7442356f1f33e227ee40a4ef9a907ea24bd0c53c356a01b44f070b

Observation abb9c9b2-1291-4586-9c3f-a45891fbcc19 · outbound

This paper cites Episodic Reinforcement Learning in Finite.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Episodic Reinforcement Learning in Finite

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.355915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:0eb1467ea36da128993988a165cf68de3ebdfa45420fd12d3af0861977a84832

Observation 2604cc15-1e64-45f0-b27d-a2c119553618 · outbound

This paper cites Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , pages =

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.399771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:6a908a8d935ca84f1569930d787b69f12f558f0975261eb620fa87147fa2de15

Observation 6865871f-6166-4178-ad00-ffc61fb2db14 · outbound

This paper cites International Conference on Machine Learning , year =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions International Conference on Machine Learning , year =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.354755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:79043aed7efc04d281821aa1fb86f6c18dd65fedffc2652af755fbf1b42dc99d

Observation 5d6dfc3b-f803-42a3-a658-c37bc2f2bb6f · outbound

This paper cites Narrowing the Gap between Adversarial and Stochastic.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Narrowing the Gap between Adversarial and Stochastic

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.368877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b92a112a74631be1933db5f3725ea4a88ffba3a21ebab431d9fc88eb64ac77ce

Observation 40ba6c59-3631-4677-ba57-2d8ab9a7cc61 · outbound

This paper cites Proceedings of the 32nd International Conference on Machine Learning , pages =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Proceedings of the 32nd International Conference on Machine Learning , pages =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.318135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:b758caf6283ee00d03bdee70027378ca1e82206a6d3b57604a8b285436e56f0a

Observation 1283848a-a0b0-449f-91e3-da4a95ad938c · outbound

This paper cites Advances in Neural Information Processing Systems , publisher =.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Advances in Neural Information Processing Systems , publisher =

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.343968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:25f4e9c8451a931871c7a89144047a20b736f3618e6d6c6d162272232a9243a0

Observation 65a0f318-5b1e-4248-81bb-e8d6426791b9 · outbound

This paper cites Fine-Grained Gap-Dependent Bounds for Tabular.

Policy Optimization Achieves Data-Dependent Regret Bounds in MDPs with Unknown Transitions Fine-Grained Gap-Dependent Bounds for Tabular

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:52:40.383860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T06:55:27.313461Z digest=sha256:d9dcd24439332f3c404779f5fd30f02bc92c241ad9fe71ea4f966cda9ffb382a

Pith citing papers

No inbound Pith citation observations are available.