Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:54:12.780877Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.07040.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:54:12.780877Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ef70f0d3-4b29-48c4-9680-c2a66502970f · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9965448c-ccfe-4d5c-8797-9056a40d7186 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Bounded semigroups of matrices
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ef97421-6307-40dd-bbb8-6844ed0b88fa · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Single sample path-based optimization of markov chains
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 869a3cd3-5da4-42c8-8d7b-7d5f54645349 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of distributionally robust average-reward reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e2f1a7-d41a-4e96-be53-613e8558d664 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with W asserstein distance
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67561d46-dd68-4551-9c2d-42cc6c7d39d6 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Distributionally robust stochastic optimization with wasserstein distance
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 825f67de-894c-4c02-bc88-01753b950a87 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sim2real in robotics and automation: Applications and challenges
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f309759-8e64-451a-a29d-50e69cfd7db2 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A robust version of the probability ratio test
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec7c81e9-8789-41b2-8b25-686b96860ea4 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust dynamic programming
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19f8b178-7a01-43f6-9858-f64fdbe6a7ff · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Is q-learning provably efficient? Advances in neural information processing systems , 31, 2018
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c88f8ffa-82e1-41a6-9ab9-6bcc34a47707 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Learning robust policy against disturbance in transition dynamics via state-conservative policy optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c16151e2-04d4-42a4-a6bf-94912749d28a · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient for rectangular robust markov decision processes
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b2800b2-64ec-4e46-b8c4-4e056478a7aa · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning First-order Policy Optimization for Robust Markov Decision Process
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf65130-af37-4ad7-840a-5349010729bd · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Reinforcement learning in robust markov decision processes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62b951b7-0427-4e8c-8d7e-707f1a52bc06 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robustness in markov decision problems with uncertain transition matrices
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fccc989e-2617-4fa9-b37f-b38b700b2ac1 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71dc82ed-7b97-47dd-8580-1cefb697f5b9 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy optimization for robust average reward mdps
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d18695a6-251c-4982-8605-2b4a842c01f6 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning u nderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, J \
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5de33fd-8f62-47de-b381-294e48cdd827 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93319adb-9ce9-47a1-b0a8-f6ca9d1f2e87 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning A finite sample complexity bound for distributionally robust q-learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97c97d62-421b-4f4c-ae57-4b937c851a91 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Sample complexity of variance-reduced distributionally robust q-learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 734d1291-44f4-4525-afbe-f989b6af7e69 · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward markov decision processes
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1942abc-4d0c-4301-9d6a-10be62db657d · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Robust average-reward reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e0f1919-a5d5-4225-89fe-f4557761741c · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free robust average-reward reinforcement learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 126f3854-5a0f-4100-be64-f9c5f82dfeac · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Policy gradient method for robust reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06666b7c-fa97-4d1e-8f68-9f7089ee3c9a · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8431c331-d2a7-44ed-b314-063e4a2c122d · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Finite-sample analysis of policy evaluation for robust average reward reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf156b68-91e8-421d-837f-e06af6c5b5ae · outbound
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Natural actor-critic for robust reinforcement learning with function approximation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.