Pith. sign in

Paper Citation Record · LEDGER

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

As of 21 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 2 inbound Pith citation observations for arXiv:2505.12462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12462 v3

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.572614Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:15.232280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:54:13.187637Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact12
  • verified fuzzy36
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d26834a-d4ce-4fda-9441-52e06de6fbd9 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Mastering the game of go with deep neural networks and tree search.nature, 529(7587):484–489, 2016

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.012652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.012652Z digest=sha256:21299c88719be1f1af8e43c83018cea6025b2f4570a4afb078e6b556f791a655

Observation 1eb46f3e-0760-4950-887f-39c805aa1223 · outbound

This paper cites Douzero: Mastering doudizhu with self-play deep reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Douzero: Mastering doudizhu with self-play deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.021461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.021461Z digest=sha256:3439a8ce4e1fdf66dfbcca5b748285bde509e9a9692b5394f15f6c55b55fdf32

Observation b707fc15-9dc1-4da5-97b7-f4f2f50d5602 · outbound

This paper cites Honor of kings arena: an environment for generalization in competitive reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Honor of kings arena: an environment for generalization in competitive reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.026944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.026944Z digest=sha256:63b4e3dfe2e2707d60dc450bbac9af0a524aa0562690e1d891e3a222d4f4a855

Observation c806d400-8e03-4809-81c7-e5ed855a8f02 · outbound

This paper cites On efficient reinforcement learning for full-length game of starcraft ii.Journal of Artificial Intelligence Research, 75:213–260, 2022.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On efficient reinforcement learning for full-length game of starcraft ii.Journal of Artificial Intelligence Research, 75:213–260, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.032504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.032504Z digest=sha256:94b8c16d60f8586cb3d607593e554fb6bd7c95379988fb476d75f033e4bba12b

Observation e3a664c7-e609-4c0f-baaf-8cd19f03444b · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.038467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.038467Z digest=sha256:568be788356596537b9d29fb020a7ccf9dbd1b05f8a8e7c0df3b92b9cc734d31

Observation 3fe4120d-b114-4d0f-8c58-fd72d0565cf0 · outbound

This paper cites Sim-to-real transfer of robotic control with dynamics randomization.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sim-to-real transfer of robotic control with dynamics randomization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.043424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.043424Z digest=sha256:d3abb2d28f41fa75ecdf7c70afdbd2798f2d52c17c53215352e923ab4669afd5

Observation f7885c32-810a-4504-8718-63b1ba83a855 · outbound

This paper cites Domain randomiza- tion for transferring deep neural networks from simulation to the real world.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Domain randomiza- tion for transferring deep neural networks from simulation to the real world

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.050564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.050564Z digest=sha256:8fdeebe8e498f6aebcdbda7f874f2bcb19c019a10dcbf10d9ba677af5ec102b4

Observation 4ac9a28b-e264-41ae-81b6-f199731c73b6 · outbound

This paper cites Deep reinforce- ment learning that matters.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforce- ment learning that matters.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.056318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.056318Z digest=sha256:c1c385d6f05667be55526394d7eb575938d4816a95710228771e1d1aa65c7632

Observation 25bf0507-06e5-4627-80ed-fd0c9b82d074 · outbound

This paper cites EPOpt: Learning Robust Neural Network Policies Using Model Ensembles.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis EPOpt: Learning Robust Neural Network Policies Using Model Ensembles

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.061975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.061975Z digest=sha256:fbeb1a10a08eb166ef3e07efb989f519aa8ca573a08bbec915fccc79e111f59c

Observation bd1772ca-b9d8-4a2d-a485-846850f30d09 · outbound

This paper cites A Study on Overfitting in Deep Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A Study on Overfitting in Deep Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.068985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.068985Z digest=sha256:dc8b508992aeff45f3a337ff25170ed8320514238f7479f27b4c3a72ec3dc083

Observation 08dfd164-3590-4ee2-aa1a-7d2bafadf905 · outbound

This paper cites Solving uncertain Markov decision processes.Carnegie Mellon University, Technical Report, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving uncertain Markov decision processes.Carnegie Mellon University, Technical Report, 2001

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.074365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.074365Z digest=sha256:b020de518b27eeef4339ccf1d50aaa3740e5cffe4f53a7c26c656a0d0af73a37

Observation 88ad3ab9-4db6-42cf-8c00-31354b7910e2 · outbound

This paper cites Robustness in Markov decision problems with uncertain transition matrices.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robustness in Markov decision problems with uncertain transition matrices

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.079730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.079730Z digest=sha256:e45914fc2d9bd9e70fd2d1321b3f6e226e05168f217d9797f0ba7db0fcf7e71f

Observation c2e870ad-258f-4c4e-8d40-b7002998c790 · outbound

This paper cites Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust dynamic programming.Mathematics of Operations Research, 30(2):257–280, 2005

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.085098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.085098Z digest=sha256:2e0dcfaaf50ed2fc8f034cc75ef17e725c4a44922ff6a1e71245f63f31f494c5

Observation 9a6c6295-57b6-4160-849f-4f7d373254a0 · outbound

This paper cites Robust adversarial reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust adversarial reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.091133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.091133Z digest=sha256:d7862da588e83d28876978ac92c88e86fb8e25bf15d96d6ac858d3b5f09dfc9c

Observation 4182383a-6e32-4580-9a9c-2b84f367c332 · outbound

This paper cites Atia, and Yue Wang.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Atia, and Yue Wang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.096890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.096890Z digest=sha256:24edc7c7878162042b590453b22d3426cd0a725d1c2034e5faf4ff44ba3b4182

Observation 8e3e992f-0cc6-48ed-bf4e-7d7b9e3a16c6 · outbound

This paper cites A reinforcement learning method for maximizing undiscounted rewards.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reinforcement learning method for maximizing undiscounted rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.102249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.102249Z digest=sha256:c05300c1cd6e34a25dc8f5aad1c9e8da51fdf5629bedb5053bed955b82c863d5

Observation 48519a1a-69e1-4789-9419-355483dd2feb · outbound

This paper cites True online td (lambda).

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis True online td (lambda)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.108024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.108024Z digest=sha256:fe85b10ff33d579aeb4d5d4f96175dde335b11810b5f4a18f4c562f7702e0666

Observation 06b8b386-9eff-4e76-9fe0-db9c8775cd1c · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.114297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.114297Z digest=sha256:2e60d0c400da0db3f20640405acb3c75eb192c3b1e82c4619f9d1fdba266916f

Observation 10e188f0-2da0-495c-930f-e6ac0e546bfc · outbound

This paper cites Learning algorithms for Markov decision processes with average cost.SIAM Journal on Control and Optimization, 40(3):681–698, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning algorithms for Markov decision processes with average cost.SIAM Journal on Control and Optimization, 40(3):681–698, 2001

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.120109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.120109Z digest=sha256:b30113b74480aca665d5f5f00435ec6c9ab3994b49e8ed50964b5245265ad7a7

Observation 5a9308df-8781-424a-b142-ecf4db2669bb · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 32(11):1238–1274, 2013

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.125399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.125399Z digest=sha256:f678c8e6e7e16eb073d7836c97ebefff0b9c9db7a7319e2ed58d878d2f74e841

Observation 8931bd48-0615-4b72-95af-ce4eed826168 · outbound

This paper cites A dynamic pricing demand response algorithm for smart grid: Reinforcement learning approach.Applied energy, 220:220–230, 2018.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A dynamic pricing demand response algorithm for smart grid: Reinforcement learning approach.Applied energy, 220:220–230, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.649236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.130372Z digest=sha256:5d074c83cd57c20139351046851997eef1fa4fee9a2d03831422ee5c9305b890

Observation 47f7954c-59c0-4c84-a761-5655afba3ce8 · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:45:57.625804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.135451Z digest=sha256:ab5d2570d5eef384c317a436bc45a5c95d62458642815013eaa6a2a6b5379d23

Observation 45aec443-a977-4933-b7e0-e2c000978f6b · outbound

This paper cites an unresolved cited work.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:45:57.603775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.140975Z digest=sha256:58d5568827f4e17fde933efe62c6207b3291df8d6b82e4f8fdfc229a20775fe5

Observation 46338669-06e2-477f-b2b1-171aa9535c93 · outbound

This paper cites Learning to trade via direct reinforcement.IEEE transactions on neural Networks, 12(4):875–889, 2001.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning to trade via direct reinforcement.IEEE transactions on neural Networks, 12(4):875–889, 2001

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.580920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.147357Z digest=sha256:a0573b97c18877d3ec0dbc80bb5eda1524c9221b53d109b4112921e177102179

Observation 0d14b53b-01d8-40e8-a516-9178f6e221de · outbound

This paper cites Reinforcement learning in economics and finance.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in economics and finance

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.560183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.153980Z digest=sha256:e4362e02b436b0172a47ca0b861f803721e5c18e8adea5fbe3904e9972926aac

Observation 72157e8a-bd7e-4818-ba9f-e85f85f0fa9a · outbound

This paper cites PhD thesis, Université d’Ottawa/University of Ottawa, 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis PhD thesis, Université d’Ottawa/University of Ottawa, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.538739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.160356Z digest=sha256:e7538ffe1ee588c8ca4f6a3412b9c6412311ac25af050e925d268a354a9415e1

Observation 3fe792d5-1114-4f2e-ac90-bd4b1d151b1b · outbound

This paper cites Deep reinforcement learning model for stock portfolio management based on data fusion.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Deep reinforcement learning model for stock portfolio management based on data fusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.516416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.166554Z digest=sha256:6f81ecd7a659ee9a4c649b6ec3fd02cebdc2593faed10e51c31a2db5a40decf1

Observation c504f33f-4f56-49b2-9216-1206f387d8e1 · outbound

This paper cites John Wiley & Sons, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2013

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.171757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.171757Z digest=sha256:64d3e74c1b59b3ab3fda2c86668712eba72f8422f6982820d39fa873a984bf15

Observation 312039e8-b8a7-4799-b573-f31a2ba735b3 · outbound

This paper cites Model-free robust average-reward reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust average-reward reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.177706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.177706Z digest=sha256:3f01078caac233b9a45d9b8e5b18a2169045ae8215cfc1440a58076fa24b02d0

Observation ac0ab864-bac9-42f7-8543-14b90c043a3e · outbound

This paper cites Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Beyond discounted returns: Robust Markov decision processes with average and Blackwell optimality

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.182579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.182579Z digest=sha256:8002b76ecfad024c7b28b7bd923be3c1913e169c386cf57ffa05095c72d6bef4

Observation ff63a073-7486-48e5-a640-243d1f8a9835 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.187895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.187895Z digest=sha256:0fcfcce16760c559463138217023f20cac2746e6d7db580379d4bf7f5bde111a

Observation 03c3a3cf-0150-4f91-99a2-51666ec12ac3 · outbound

This paper cites Span-Based Optimal Sample Complexity for Average Reward MDPs.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Span-Based Optimal Sample Complexity for Average Reward MDPs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.598129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.193656Z digest=sha256:197feb1b2599edab5c03840c6cb3874a2757af1e922ab349ebd9c117ff50c5e7

Observation ebba0abb-edd2-4f8c-bd25-42c2b4be8d86 · outbound

This paper cites Robust average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust average-reward Markov decision processes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.448638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.199546Z digest=sha256:2a90854e634b1cfffaec87cfa1dc024ab5c1f062f8247ea5374f75e62c472ec6

Observation f1279985-aebc-48a8-9c57-449479ff1227 · outbound

This paper cites Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.570520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.205307Z digest=sha256:7e76cbb3651b339d8d17cef46199768796558840749e2b87d832b0e0c2b5b003

Observation 7df5871f-2f1b-4e21-9a7d-d4dc57982ba2 · outbound

This paper cites A reduction framework for distributionally robust reinforcement learning under average reward.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A reduction framework for distributionally robust reinforcement learning under average reward

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.427393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.210962Z digest=sha256:384f7f80229b693fddbf9cb2c3d8c3194989976c56d501a03ef44e148da41b13

Observation 3ef557fc-1712-4ea4-a58e-221803ccd085 · outbound

This paper cites Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.540031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.215882Z digest=sha256:705d91f663e621974812d8c01ce17d33ed78b4b37a1f00e60df92a29a79ffded

Observation 731997b1-12fc-4481-89e5-315d3905efa7 · outbound

This paper cites Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.arXiv preprint arXiv:2502.16816, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.arXiv preprint arXiv:2502.16816, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.221589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.221589Z digest=sha256:42ae335afe759f2e89e4c2c3668638416b0f07239ebbb06c4f752153f832d461

Observation e9fbe852-c90f-46ad-be6b-e15f0270eb98 · outbound

This paper cites Sample complexity of distributionally robust average-reward reinforce- ment learning.arXiv preprint arXiv:2505.10007, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust average-reward reinforce- ment learning.arXiv preprint arXiv:2505.10007, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.226946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.226946Z digest=sha256:546829c43910a41665c8ba427a82de620aff7bd3aac0b3f02383581472dfc2da

Observation 689c66a9-3d39-42a9-8118-180e9d37ad66 · outbound

This paper cites Dynamic Programming and Optimal Control 3rd edition, volume II.Belmont, MA: Athena Scientific, 2011.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Dynamic Programming and Optimal Control 3rd edition, volume II.Belmont, MA: Athena Scientific, 2011

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.403509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.233628Z digest=sha256:b277f7f332a48e76f31a78a1fe20de58fa1fe43a7ebded79b57c845b2098fe96

Observation 3bf91cab-7322-4085-89b1-f779c1e5d78b · outbound

This paper cites Fixed points of nonexpanding maps.Bulletin of the American Mathematical Society, 73(6):957–961, 1967.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Fixed points of nonexpanding maps.Bulletin of the American Mathematical Society, 73(6):957–961, 1967

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.374939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.239150Z digest=sha256:57e1cd8f80d30dec20736ed9db1cb8510470037772edf7ffcfb1568322892883

Observation 231e7c11-9dcc-4e25-b2fa-85442057c34b · outbound

This paper cites On the convergence rate of the Halpern-iteration.Optimization Letters, 15(2):405–418, 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the convergence rate of the Halpern-iteration.Optimization Letters, 15(2):405–418, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.351750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.244118Z digest=sha256:5afc504bf9e1db578555bf2f3dba9a24cfffa88ca72731c96650d6e07721fdfd

Observation f6508e46-19b5-4062-942a-46c20f72112f · outbound

This paper cites Near-Optimal Sample Complexity for MDPs via Anchoring.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near-Optimal Sample Complexity for MDPs via Anchoring

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.372929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.249423Z digest=sha256:c61b49fbec3c04b48c37f3dbc3ec5abb03b906cc298ff5d236b9c9ea2ea687f9

Observation 9c944db6-5aa7-41bb-a437-86e8f9e5fdf7 · outbound

This paper cites Online robust reinforcement learning with model uncertainty.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Online robust reinforcement learning with model uncertainty

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.327574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.256185Z digest=sha256:9b7be94ea33f427006d4c007a9bf23877997a16e6f885e91e2c3eef977877b50

Observation 79954e55-b9b0-4351-b1de-d7a892a9c78a · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Policy gradient method for robust reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.305017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.261275Z digest=sha256:fcb1d5188f3d88ab442d377fd69017aa1dec95c8b7192b9ef925131f36d36493

Observation 4957033f-4b06-4e80-8eec-b8f1f563a3f2 · outbound

This paper cites Minimax-Optimal Multi-Agent Robust Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax-Optimal Multi-Agent Robust Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.347346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.267200Z digest=sha256:face4f7a5ec08a7c2a948bf6973be4112857b29cb91c8c160332a6f8bc62a8e5

Observation a70223d0-d33b-4d53-ab1a-7619e15e51a2 · outbound

This paper cites An Efficient Solution to s-Rectangular Robust Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis An Efficient Solution to s-Rectangular Robust Markov Decision Processes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.273305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.273305Z digest=sha256:5b112e0ab6aeacda2ceb8740fb32095941d18ef77a6ab45457161d781fe1b644

Observation a38a16e3-d437-45c0-b66f-ce924df465ff · outbound

This paper cites Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183, 2013

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.282637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.278368Z digest=sha256:e379338f8c42637496a2093164d22a594fadd950ad28d87502231bab48b4c231

Observation 86f4ac45-2605-42ef-a246-ae1e404bc3ef · outbound

This paper cites Sample complexity of robust reinforcement learning with a generative model.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of robust reinforcement learning with a generative model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.283480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.283480Z digest=sha256:21f6285ff1051ba3a246182f0ae4f3ef7658d5f3df29cd2961f6fff8dd0fcc6e

Observation b770c689-eddb-4ee5-8fe2-74dabab080e7 · outbound

This paper cites The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.288950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.288950Z digest=sha256:9ba3843cbbbfaeb98df89e0dbff53d56fe9257cd9b710d0d54cdd3121bdbf039

Observation 73e592bd-9389-4697-843d-8bc9fcd08f98 · outbound

This paper cites Improved sample complexity bounds for distributionally robust reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Improved sample complexity bounds for distributionally robust reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.244736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.294763Z digest=sha256:29816b4644d9600b0d106f0eb2c0033769039b8f5cd4cfb9964597be58d29e78

Observation 2aedd444-f311-4d75-82a7-65b9c1fc94ec · outbound

This paper cites Learning and planning in average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Learning and planning in average-reward Markov decision processes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.223307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.300280Z digest=sha256:97a38ba766ef2c7333e0e56464e6022e6a3a60876e7f40fe80683f967862f181

Observation f0c00779-a547-44fc-9bcb-341c24aade4a · outbound

This paper cites On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On Convergence of Average-Reward Off-Policy Control Algorithms in Weakly Communicating MDPs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.305675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.305675Z digest=sha256:d3b659bcf5590314be8ff0fda99d67f35e3d6a03c664721302235a5036849a8f

Observation 368010cc-a2a8-46d9-92c0-8c2b32488ac2 · outbound

This paper cites The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.311465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.311465Z digest=sha256:6aa6e9b48712d53ae166a211fd7a9400a7f2cc64f992446c5eae5a2b4553321e

Observation 4fc46e46-8a10-4048-8159-a798b104e511 · outbound

This paper cites Sharper model-free reinforcement learning for average-reward Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sharper model-free reinforcement learning for average-reward Markov decision processes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.200349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.318155Z digest=sha256:23b4a9ebc157c06e9318d8a2898e4653cbbe9ef9ecb4a18076d9b5b47b5cb9e8

Observation 338b6fba-9057-4f73-81b2-e094b64a5a5f · outbound

This paper cites Finite sample analysis of average-reward TD learning andQ-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite sample analysis of average-reward TD learning andQ-learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.180856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.323713Z digest=sha256:3e2e5d77cb356af0577773b129673b9055b7bd0ee2f18b726cb53dc0c0462d00

Observation 16119503-5435-46f9-b367-31935e38d1b4 · outbound

This paper cites A first order method for solving convex bilevel optimization problems.SIAM Journal on Optimization, 27(2):640–660, 2017.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A first order method for solving convex bilevel optimization problems.SIAM Journal on Optimization, 27(2):640–660, 2017

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.330060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.330060Z digest=sha256:0b2741aabde3141f01444eaee2370136a1e85f9bcc6546450e60a79101a86750

Observation 8281b926-a697-47fe-9c30-0197123579d3 · outbound

This paper cites Exact optimal accelerated complexity for fixed-point iterations.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Exact optimal accelerated complexity for fixed-point iterations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.147868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.335227Z digest=sha256:0fe05c2ef6f43302a3e7ae22907d8c0fde7c7d0eae220ba2d24ab0135f08eef5

Observation a810d212-e376-4b4a-bd04-2c33a966d48e · outbound

This paper cites Optimal error bounds for non-expansive fixed-point iterations in normed spaces.Mathematical Programming, 199(1):343–374, 2023.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal error bounds for non-expansive fixed-point iterations in normed spaces.Mathematical Programming, 199(1):343–374, 2023

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.341859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.341859Z digest=sha256:bdfbeb41843fbbe8f3c8edefe56fc5fe5e33218c25325acb831fb03878d12071

Observation 7e7f076f-034e-4a4a-bd36-18af452e264c · outbound

This paper cites Optimal non-asymptotic rates of value iteration for average-reward markov decision processes.arXiv preprint arXiv:2504.09913, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal non-asymptotic rates of value iteration for average-reward markov decision processes.arXiv preprint arXiv:2504.09913, 2025

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:56.251252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.346894Z digest=sha256:5037ada0e71784da50e70e9343fec6dbc5592dc3562e9f1fb44867d1b98804c0

Observation 0179c015-5c57-402a-8a54-9d2524c23b05 · outbound

This paper cites Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.352235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.352235Z digest=sha256:b43e3d5996b54290250aec3745b3f9bb286b8bfa154f5cce6bd1fa46b3258cc8

Observation 73ea37a3-fa74-4ac4-ad76-62c9d6d337a5 · outbound

This paper cites Minimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning.arXiv preprint arXiv:2403.09621, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Minimax optimal and computationally efficient algorithms for distributionally robust offline reinforcement learning.arXiv preprint arXiv:2403.09621, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.358774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.358774Z digest=sha256:23564d14f4febae7fa1dd6244469bee9c8745870a9d1eea56d9579527519af92

Observation 41b163ad-42a6-4882-bdf9-0dca89716d64 · outbound

This paper cites McGill University (Canada), 2021.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis McGill University (Canada), 2021

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.107802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.365701Z digest=sha256:84756509d88c45d5254cb873e72ac56d29d25ff43539c3af29f53fe86898060c

Observation 5556fe4f-b56f-4df7-8f14-823e3d65a13a · outbound

This paper cites Distribution- ally robustQ-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distribution- ally robustQ-learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.085211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.370858Z digest=sha256:8a1dc4120e0fac81a30d72fb5ba770285f9da36619d3f8894b9f6e18e91f5a41

Observation a6a48b0e-fe83-437e-a98e-c9454632999b · outbound

This paper cites Truncated Variance Reduced Value Iteration.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Truncated Variance Reduced Value Iteration

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.376217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.376217Z digest=sha256:c60f53fb7e5a5f4925d1b7597f94f3ca5c1c7e6721205524e0e170012f929918

Observation afefc198-8f6d-4aa6-af39-c1728aa97aae · outbound

This paper cites Optimal approximation of average reward markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal approximation of average reward markov decision processes

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.061174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.383742Z digest=sha256:ceac8d8f9225d49c38dd6b0fba9e39bb9a0913d62bfa0bc6420b5c755f55872f

Observation 6291b4b6-4fc3-4311-bbb2-2b5e064cb378 · outbound

This paper cites Optimal Sample Complexity for Average Reward Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity for Average Reward Markov Decision Processes

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.060257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.389158Z digest=sha256:391c5e9c44f4fa42d8118f1d0e65ca28ebd4389dfa694a0778b00913f950a6e2

Observation bbcd3ccd-dc07-4cd4-a930-e02b95d2e698 · outbound

This paper cites Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Optimal Sample Complexity of Reinforcement Learning for Mixing Discounted Markov Decision Processes

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:56.034715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.394758Z digest=sha256:ee35456cd0dba5960b9c0df58b09e385b0505df5df77c7a47561bc5b708ba58a

Observation ec06317d-4efe-436b-9f2a-14d16d134575 · outbound

This paper cites Feasible q-learning for average reward reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Feasible q-learning for average reward reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.037853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.400232Z digest=sha256:e3f17c60f42ab4d4262e96de70bc844ccd240ae7c1ff07c5860d19b2941caed6

Observation f8309385-1e25-41c3-b22f-454f64601c69 · outbound

This paper cites Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.405318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.405318Z digest=sha256:fc3058a43edc74b8b79ddcf47803a1548e0d7962546dab217f0f96079bbc7557

Observation d8ccf8a8-f416-4469-8e24-c1d9a1b5e8ad · outbound

This paper cites Tightening the dependence on horizon in the sample complexity of q-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Tightening the dependence on horizon in the sample complexity of q-learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:57.017859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.411017Z digest=sha256:8b20f630619f7679b122e3433d1841a8a87b5d43e3c60d46cdb028ac1abd4cf9

Observation 6d3af054-8da4-48b8-95a8-8e61ef415611 · outbound

This paper cites Achieving the asymptotically minimax optimal sample complexity of offline reinforcement learning: A DRO-based approach.Transactions on Machine Learning Research, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Achieving the asymptotically minimax optimal sample complexity of offline reinforcement learning: A DRO-based approach.Transactions on Machine Learning Research, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.998296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.416316Z digest=sha256:41859f87bee1efe115046bf136e01dd50c9b6f7332577948bf697e675cb611a7

Observation 3475e849-e554-4f88-a71e-ee3ac9f89f61 · outbound

This paper cites Gambling in a rigged casino: The adversarial multi-armed bandit problem.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Gambling in a rigged casino: The adversarial multi-armed bandit problem

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.421904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.421904Z digest=sha256:a71fd1d12511d27f6b7ec2b727ac0adcf52939ffda7a92fae586821a348bbadf

Observation fbc64807-e11b-4f17-a8a1-5a588ff2601b · outbound

This paper cites What Doubling Tricks Can and Can't Do for Multi-Armed Bandits.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis What Doubling Tricks Can and Can't Do for Multi-Armed Bandits

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.427082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.427082Z digest=sha256:318b2c9cb41a54e433f40cfa622f3c285ff40a01beadf7494aa7db2dbea0e1dc

Observation d620034c-9e67-4f37-a518-62646140225c · outbound

This paper cites Reducing blackwell and average optimality to discounted MDPs via the blackwell discount factor.Advances in Neural Information Processing Systems, 36, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reducing blackwell and average optimality to discounted MDPs via the blackwell discount factor.Advances in Neural Information Processing Systems, 36, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.965390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.432919Z digest=sha256:81a7d3e3bdb1bce0b1b0e5d673b503005dc515d975ee2a78b9fdcc06f88b9ca5

Observation 3643a70c-a5ad-47f9-a32a-602042336f2d · outbound

This paper cites Finding good policies in average-reward Markov Decision Processes without prior knowledge.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finding good policies in average-reward Markov Decision Processes without prior knowledge

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:55.971177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.437973Z digest=sha256:071ac44690fff8a01f7ad7ab3d0f1bcc8c976ab73fc8fafd82218931a20e23ae

Observation 797ef007-ed44-46d6-93e1-003df84b07f1 · outbound

This paper cites Model-free robust reinforcement learning with sample complexity analysis.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Model-free robust reinforcement learning with sample complexity analysis

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.947229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.443468Z digest=sha256:d01c6e9d175c8accb73a3cca940e1b709fcaa617d9707ccdcdae6ad6cc01b728

Observation bef91495-4b79-44ee-b2d2-151a0645a5ed · outbound

This paper cites Bounded parameter Markov decision processes with average reward criterion.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bounded parameter Markov decision processes with average reward criterion

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.924798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.449685Z digest=sha256:fbce3fe7588a2b8887ab5ca74375e39c8322fba20da287b0d5cafa24278a8c4a

Observation fcd4fd70-0b96-4efa-b7ef-a3b957949251 · outbound

This paper cites Bellman optimality of average-reward robust markov decision processes with a constant gain.arXiv preprint arXiv:2509.14203, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bellman optimality of average-reward robust markov decision processes with a constant gain.arXiv preprint arXiv:2509.14203, 2025

Reference 78

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:55.945028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.456436Z digest=sha256:4a8fc4f53083d63bcba59347f36a134189367889ac262d4c4183e7f4e54e60d7

Observation cf7f92d9-6f21-470b-8ce1-f1505ebfcdd4 · outbound

This paper cites Solving Long-run Average Reward Robust MDPs via Stochastic Games.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Solving Long-run Average Reward Robust MDPs via Stochastic Games

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.462391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.462391Z digest=sha256:3dc3a563005c22317da3a3ae61d94f6dfa2210ce7e7e8b52f5b412b0219c0c84

Observation 0323bc47-76a1-48a6-ab37-ffcac4ab488e · outbound

This paper cites Reinforcement learning in robust Markov decision processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Reinforcement learning in robust Markov decision processes

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.902355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.470239Z digest=sha256:21baf059f4c5405d33ba317f1d4ac1829f107a4d62c36e94ab7941134b301ea6

Observation 83e1f644-ffc9-4576-b919-8d896900f98c · outbound

This paper cites Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics.The Annals of Statistics, 50(6):3223–3248, 2022

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.476517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.476517Z digest=sha256:7567f5810df093e650d675b5e4515209519f9cc191b6ddebae7086114a72e034

Observation 288d8550-4d1c-473e-a74e-e10ee77594f4 · outbound

This paper cites Finite-sample regret bound for distributionally robust offline tabular reinforcement learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Finite-sample regret bound for distributionally robust offline tabular reinforcement learning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.871560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.483191Z digest=sha256:724388e00d3ea2c2f555bd9a5188f0d40bcb169bd5b341d402501ed5fb0d93fc

Observation b5d591de-c93f-4a01-835b-3f9633b1cb0e · outbound

This paper cites A finite sample complexity bound for distributionally robustq-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A finite sample complexity bound for distributionally robustq-learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.851425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.488217Z digest=sha256:fbdc0dbe7753909b4e7225c78ad0290561e1253c40eaef2d415c9a8057c45408

Observation ae1f8c9d-162f-4da7-a650-84c95a131da9 · outbound

This paper cites Single-Trajectory Distributionally Robust Reinforcement Learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Single-Trajectory Distributionally Robust Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.493828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.493828Z digest=sha256:fbfd0b1c0f9219ee9fb93c2d6c6bae8b13e431b6778d3ab4a5d1db40396bf32e

Observation 64d8134c-cc7e-45f0-9658-1ea0d8ee2f54 · outbound

This paper cites Sample Complexity of Variance-reduced Distributionally Robust Q-learning.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Variance-reduced Distributionally Robust Q-learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:45:55.821668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.499344Z digest=sha256:e2217e42e35a7173d12a23bc0703b09c1604efe4f40306eff42cc49d984d0bce

Observation bb58ba8b-5994-49a0-b605-3b8d7e2c0dc6 · outbound

This paper cites Bring your own (non-robust) algorithm to solve robust mdps by estimating the worst kernel.arXiv e-prints, pages arXiv–2306, 2023.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Bring your own (non-robust) algorithm to solve robust mdps by estimating the worst kernel.arXiv e-prints, pages arXiv–2306, 2023

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.831060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.504737Z digest=sha256:aa9754593e95a3eb2ae394c93f5ea2a29130da83b12ede5d97976d3b7de69bf7

Observation b1595b8c-b9ef-4564-acbc-dafc5387a8dc · outbound

This paper cites Twice regularized MDPs and the equivalence between robustness and regularization.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Twice regularized MDPs and the equivalence between robustness and regularization

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.812885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.509603Z digest=sha256:983bc12b9a4a64c14476db49d33716bdde7406c4a51fb10d2d51494033a91a4b

Observation 10f49b19-4ad2-4f2e-b373-bfebbe1e3168 · outbound

This paper cites Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.514519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.514519Z digest=sha256:981df9d031b384e6a1a2eee656d55b48d7cac4fec9a16ab6ccbf0ee112c05226

Observation bf25c8fe-24d5-4a71-b7e8-fee6f65a26e0 · outbound

This paper cites Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.519507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.519507Z digest=sha256:652cfc1e1ba004ef79c643a47f9609715fb31c312afaa7036370a0e31c1c1544

Observation 44a08380-d943-469e-81c6-e6543d0afad0 · outbound

This paper cites A unified principle of pessimism for offline reinforcement learning under model mismatch.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis A unified principle of pessimism for offline reinforcement learning under model mismatch

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.793150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.525300Z digest=sha256:c29ed802af3f919cea93e360f39d028bceb06919c6a3847a2433e04cb0763b0b

Observation c237809b-e432-4ccc-8c39-684560ff0331 · outbound

This paper cites Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.531036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.531036Z digest=sha256:6044419fb649b3f0a6d51643b08817bbcc515e4376a3f2ca1257bc11da3a7b63

Observation 8004d090-d143-47e0-b57e-770675366687 · outbound

This paper cites Provably near-optimal distributionally robust reinforcement learning in online settings.arXiv preprint arXiv:2508.03768, 2025.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Provably near-optimal distributionally robust reinforcement learning in online settings.arXiv preprint arXiv:2508.03768, 2025

Reference 92

Resolution
verified exact
raw_fallback, observed 2026-08-15T20:45:55.736471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.536974Z digest=sha256:f9b601e6c2e8a8c35c17fcce3a94dbef067ae0dd0bb44d8cec3b028c21456914

Observation 1ab2fa1c-d692-4b1d-8c1a-b1ed925655ca · outbound

This paper cites Sample complexity of distributionally robust off-dynamics reinforcement learning with online interaction.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Sample complexity of distributionally robust off-dynamics reinforcement learning with online interaction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.542380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.542380Z digest=sha256:75fff768c92a41e173df2cc75c847f6e4f82d16ccfdc76102a61511b522b4fe6

Observation 3913a34b-a13e-439b-af82-cd7f19adc41b · outbound

This paper cites John Wiley & Sons, 2014.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis John Wiley & Sons, 2014

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.548339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.548339Z digest=sha256:850bbbbc7ea7680fb7e7c19d36748cbdaa5027a1931e6e91186372742869bb52

Observation 934d4e26-d97e-477c-85f3-127962baae9c · outbound

This paper cites Average-reward model-free reinforcement learning: a systematic review and literature mapping.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Average-reward model-free reinforcement learning: a systematic review and literature mapping

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.553994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.553994Z digest=sha256:cf3f88ba7484804d3e0bd0851722145bec459fb9023d6b1b9d997636d3a0656b

Observation 929fa1d8-5f50-4949-b9d1-fdd170f0395c · outbound

This paper cites Towards tight bounds on the sample complexity of average-reward mdps.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Towards tight bounds on the sample complexity of average-reward mdps

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.746953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.560171Z digest=sha256:4f2c6db1d1040e0846a97a7b1738c54fa44078dcfcd312c0b365c7e51f7df2d4

Observation 98370f72-38ee-441f-8184-fbd7b96c6cf2 · outbound

This paper cites Stochastic first-order methods for average-reward markov decision processes.Mathematics of Operations Research, 2024.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Stochastic first-order methods for average-reward markov decision processes.Mathematics of Operations Research, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.726923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.565897Z digest=sha256:128b63e6772640fee7213d4e0a4e229a0819b461938276712092651b7ecdfc92

Observation 518ee3bf-f958-46b1-bcba-8b3836481063 · outbound

This paper cites On the generation of Markov decision processes.Journal of the Operational Research Society, 46(3):354–361, 1995.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis On the generation of Markov decision processes.Journal of the Operational Research Society, 46(3):354–361, 1995

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:45:56.706015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T20:45:55.572614Z digest=sha256:896e403465440967caaa3df643e10948495c054a47abdd610ee7cf729c49c446

Pith citing papers

Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:54:13.272094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.619397Z digest=sha256:c9251ca249da7a665f8d1340d4f2f3cb24ff7e8ecaf89240901980ff171a3a8c

Observation d541f5cd-ef54-46b1-8f06-ffb1343a118f · inbound

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions cites this paper.

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:15.232280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:15.232280Z digest=sha256:950174b59f409a0d5342f725ff1ece6def36186cb95067a3f9664e83fc715362