Pith. sign in

Paper Citation Record · LEDGER

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.07040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07040 v4

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:54:12.780877Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef70f0d3-4b29-48c4-9680-c2a66502970f · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning On the theory of policy gradient methods: Optimality, approximation, and distribution shift

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.474894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.616419Z digest=sha256:6ab2e05b9ea7ce2ff496541f5eb48d26569ca2c0664ef0672f7cb51485ae069c

Observation 9965448c-ccfe-4d5c-8797-9056a40d7186 · outbound

This paper cites Bounded semigroups of matrices.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Bounded semigroups of matrices

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.297043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.707205Z digest=sha256:c59c48d8de73ae8a5bb4c341873e8d78f7d13c6e08ed083ec0d414248035081b

Observation 9ef97421-6307-40dd-bbb8-6844ed0b88fa · outbound

This paper cites Single sample path-based optimization of markov chains.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Single sample path-based optimization of markov chains

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:17.111742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:09.869727Z digest=sha256:7f8912900f3dafdf56a389e244e65293bf43876c75dac456bcddd1786a8ca6d2

Observation 869a3cd3-5da4-42c8-8d7b-7d5f54645349 · outbound

This paper cites Sample complexity of distributionally robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of distributionally robust average-reward reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:09.997942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:09.997942Z digest=sha256:080e8650b7ee07b70081e1dc5292e5037768807724eed0c2d8fc52be253f1884

Observation 61e2f1a7-d41a-4e96-be53-613e8558d664 · outbound

This paper cites Distributionally robust stochastic optimization with W asserstein distance.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with W asserstein distance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.886561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.170591Z digest=sha256:015f5a7c8ed0b4cd0c9c0f1f42782c253f139eea6e99f322584fe584169b7843

Observation 67561d46-dd68-4551-9c2d-42cc6c7d39d6 · outbound

This paper cites Distributionally robust stochastic optimization with wasserstein distance.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with wasserstein distance

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.737358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.326835Z digest=sha256:897c3db9a1f447bde8031f3717203865b78afb148abe01b568b44f9acde6d064

Observation 825f67de-894c-4c02-bc88-01753b950a87 · outbound

This paper cites Sim2real in robotics and automation: Applications and challenges.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sim2real in robotics and automation: Applications and challenges

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.572891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.442042Z digest=sha256:ba1c6b4a0d066652805c60fa71d03d620a7a4af1e8a430e1b01867c48efb492f

Observation 7f309759-8e64-451a-a29d-50e69cfd7db2 · outbound

This paper cites A robust version of the probability ratio test.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A robust version of the probability ratio test

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.386215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.553831Z digest=sha256:b65fcf707c59321b06713576af6a701ae95ffe642b41ae7d3d64ab8a422cf495

Observation ec7c81e9-8789-41b2-8b25-686b96860ea4 · outbound

This paper cites Robust dynamic programming.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust dynamic programming

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:16.227215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.666749Z digest=sha256:bc8a52548a3b5fc2fad2dc458a926578178d469a684103471469831c789d26bb

Observation 19f8b178-7a01-43f6-9858-f64fdbe6a7ff · outbound

This paper cites Is q-learning provably efficient? Advances in neural information processing systems , 31, 2018.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Is q-learning provably efficient? Advances in neural information processing systems , 31, 2018

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.999655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.782406Z digest=sha256:5ce3f07b1ae3743df5856c4262f7d9ce1c7bfc89698245bc86892090093564e4

Observation c88f8ffa-82e1-41a6-9ab9-6bcc34a47707 · outbound

This paper cites Learning robust policy against disturbance in transition dynamics via state-conservative policy optimization.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Learning robust policy against disturbance in transition dynamics via state-conservative policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.849079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:10.965109Z digest=sha256:dc93b34b70c2451ea90bb541c64d518efb56544fcbeb46aa7d664299bcb6f6d1

Observation c16151e2-04d4-42a4-a6bf-94912749d28a · outbound

This paper cites Policy gradient for rectangular robust markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient for rectangular robust markov decision processes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.670114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.050545Z digest=sha256:b23a0e5fc8a33634446a92413f62a6482224ec0b6efe44f28c51ff6b3fac4458

Observation 6b2800b2-64ec-4e46-b8c4-4e056478a7aa · outbound

This paper cites First-order Policy Optimization for Robust Markov Decision Process.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning First-order Policy Optimization for Robust Markov Decision Process

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.213481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.213481Z digest=sha256:676d62915a7148555d835ba156c2bf0ef3d04c0c8d25d9f339e2761a66cfb50f

Observation 1bf65130-af37-4ad7-840a-5349010729bd · outbound

This paper cites Reinforcement learning in robust markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Reinforcement learning in robust markov decision processes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.458906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.344482Z digest=sha256:872d2596d50630a460f10a1ea4983f188e160bc096c7755ac5de330da2b949ed

Observation 62b951b7-0427-4e8c-8d7e-707f1a52bc06 · outbound

This paper cites Robustness in markov decision problems with uncertain transition matrices.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robustness in markov decision problems with uncertain transition matrices

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.289414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.492619Z digest=sha256:4963576345b4501a480672a24c342786874fa88f3b9e54b6cbd08e90e176a4ea

Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · outbound

This paper cites Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:54:13.272094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.619397Z digest=sha256:28090256a2d59cfcab42b9565222548132ba0f5c3fbba64a37ae107733609a7c

Observation 71dc82ed-7b97-47dd-8580-1cefb697f5b9 · outbound

This paper cites Policy optimization for robust average reward mdps.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy optimization for robust average reward mdps

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:15.126484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.789338Z digest=sha256:4b05967d597f51232e2ed3655adcd42c657bcc5444d3728d56d7bcb07bdcf618

Observation d18695a6-251c-4982-8605-2b4a842c01f6 · outbound

This paper cites u nderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, J \.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning u nderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, J \

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.913558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:11.912786Z digest=sha256:3e9b0df86c6e0df280cf40621d62b84f6fce5b262b271ea5a3469dbb896368a1

Observation a5de33fd-8f62-47de-b381-294e48cdd827 · outbound

This paper cites Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.972128Z digest=sha256:9a5cd989fd9e823bdd5de0c459a9dedfc5ad3ccd9a09edfaf3c5e94aa4c72a96

Observation 93319adb-9ce9-47a1-b0a8-f6ca9d1f2e87 · outbound

This paper cites A finite sample complexity bound for distributionally robust q-learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A finite sample complexity bound for distributionally robust q-learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.746932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.041666Z digest=sha256:1dee6d3425e2a83d5c1471c3df7d617638f38f1923eb164c8ae6d6b9ba819e73

Observation 97c97d62-421b-4f4c-ae57-4b937c851a91 · outbound

This paper cites Sample complexity of variance-reduced distributionally robust q-learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of variance-reduced distributionally robust q-learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.582387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.176391Z digest=sha256:e23026eec02d8699e5f3fb9f6485b7448055da11e4218128f66fd7b484deb2f4

Observation 734d1291-44f4-4525-afbe-f989b6af7e69 · outbound

This paper cites Robust average-reward markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward markov decision processes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.431435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.288138Z digest=sha256:299b8dc2f53c758d3113887d8852dfa2046e97ca34154027c80b01c642ba4de1

Observation b1942abc-4d0c-4301-9d6a-10be62db657d · outbound

This paper cites Robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.269153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.366541Z digest=sha256:6c0a4493e1bb7878975fda6fe0f89bcced98269d3e43e028e6755d73a56c0efd

Observation 5e0f1919-a5d5-4225-89fe-f4557761741c · outbound

This paper cites Model-free robust average-reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free robust average-reward reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:14.072906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.439077Z digest=sha256:6e7f3464c9179c265a66a4a49d134782ba8f9f58226c37f7cd822baa4bf7086b

Observation 126f3854-5a0f-4100-be64-f9c5f82dfeac · outbound

This paper cites Policy gradient method for robust reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient method for robust reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.920610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.514568Z digest=sha256:3da329528ea14786783c731a5ce8eb1a6c16fdd96f61212eb3c68e1efeb90aa1

Observation 06666b7c-fa97-4d1e-8f68-9f7089ee3c9a · outbound

This paper cites Model-free reinforcement learning in infinite-horizon average-reward markov decision processes.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free reinforcement learning in infinite-horizon average-reward markov decision processes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.733990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.593503Z digest=sha256:b2135b0f103b2867d38547f29c4fd977fab61574525e7b2e5a9edeb80130be77

Observation 8431c331-d2a7-44ed-b314-063e4a2c122d · outbound

This paper cites Finite-sample analysis of policy evaluation for robust average reward reinforcement learning.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Finite-sample analysis of policy evaluation for robust average reward reinforcement learning

Reference 27

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:54:13.058498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.692896Z digest=sha256:db1923f2eab5f0ff731d9cd216dc27411537beec45848804f09ad9abb2cc08d7

Observation bf156b68-91e8-421d-837f-e06af6c5b5ae · outbound

This paper cites Natural actor-critic for robust reinforcement learning with function approximation.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Natural actor-critic for robust reinforcement learning with function approximation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:54:13.578820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:54:12.780877Z digest=sha256:cefcd5c9db473d4aca5b1987250e35a452dd6c47bd4b71c31a0a57e11382a48a

Pith citing papers

No inbound Pith citation observations are available.