Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:31:38.795970Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.21314.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:31:38.795970Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 33b9f85e-83ba-409e-91d2-23c77c4164d8 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Optimal control of Markov processes with incomplete state information I,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbbb163e-c714-4370-a962-f4a5485fde5b · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs The optimal control of partially observable Markov processes over a finite horizon,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d77d49cb-4264-422f-b0f9-d3a53976a919 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Approximate information state for approximate planning and reinforcement learning in partially observed systems,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e999cc6a-1015-4d36-921d-d4476f78ce11 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Deep recurrent Q-learning for partially observable MDPs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 159ff1fe-54b4-4789-9574-1d09c12456cc · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Deep variational reinforcement learning for POMDPs,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df196e73-64fd-43da-aa25-4a7373d4756a · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs On Improving Deep Reinforcement Learning for POMDPs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3d8e18d-0db7-49a9-bafb-4eb16fa66946 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Memory-based deep reinforcement learning for POMDPs,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a2642f5-2397-4d12-acc7-5abeb1c96664 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9d00860-670b-49b2-8744-f1c80c843f50 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Agent-state based policies in POMDPs: Beyond belief-state MDPs,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3688d29-c638-4798-aae7-069a54b5d3aa · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning algorithm for partially observable Markov decision problems,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b638296f-62e1-47f6-945b-fb580ccef064 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f1c3ac5-5f2b-4344-9903-cbf5a49b5271 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Periodic agent-state based Q- learning for POMDPs,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ede0e294-ec60-42b2-8ef0-664e944b95b9 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning for stochastic control under general information structures and non-markovian environments,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27d0ba31-a0cc-42fd-92bd-10fb265bd4f2 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning in non-Markovian environments,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b90a2be-c0ae-463a-94f0-7b2268e34eb4 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs On actor-critic algorithms,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54d6498f-8d3a-4649-9795-36c8460fc974 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f54d08a-858c-4b4a-b90f-52c8076f127f · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5c8d233-b142-4b9c-868f-66b4c9e65581 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Under- standing the impact of entropy on policy optimization,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df1c5df4-e07a-464b-9283-9e5d00a0a50c · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Deep reinforcement learning that matters,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80329930-38b3-4885-89ff-c0b9b9139e3f · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Relative entropy policy search,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b2236a6-2014-46c0-9a35-d4773476fc40 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Trust region policy optimization,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4ace202-8d60-4877-9cce-60afb331329e · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs A unified view of entropy- regularized Markov decision processes,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8f17d5c-6661-4d4c-83be-9bea98b722b4 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs A theory of regularized Markov decision processes,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e46cf0c7-e872-4fe8-a4e1-1c1fa766e711 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Learning latent dynamics for planning from pixels,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21b2282f-babd-462a-9f4d-ed48a9b90bc4 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Solar: Deep structured representations for model-based reinforcement learning,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8683189-dbd0-49f4-8ad1-d2f855b17ca7 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Dream to control: Learning behaviors by latent imagination,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 544e148c-7e2a-44be-afec-00c55961a26b · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Bridging state and history representations: Understanding self-predictive RL,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e597574b-5bf1-44e1-9e16-55694b7ce7b8 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Entropy-regularized Point-based Value Iteration
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bf9b10-891e-4cd9-bba0-2bff1acb2561 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs DESPOT: Online POMDP planning with regularization,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e19dd16f-0ff1-430d-a5d8-04bc53bb48a3 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4468ba5-39a4-49b0-9778-870181cf0d07 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs The limits of pure exploration in POMDPs: When the observation entropy is enough,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c184507a-1e85-4b33-8a98-74a804bb044f · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d25e0a-98c9-422f-947b-135b42902d99 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Hiriart-Urruty and C
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a55d2cd3-1b4c-4363-89e4-79d9a2276489 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Differentiable dynamic programming for structured prediction and attention,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cfe5ea7-ef0b-483e-b378-ccb4976e26bd · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4de4b806-c3dd-4c6d-9232-57f89e0772f5 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning with deep energy-based policies,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 040ab679-c9b4-4c4e-a0d3-edf9b0ad9552 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs A stochastic approximation method,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ab540f2-df98-4d98-92f6-2610b567d596 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 119f1e27-9cfe-4cb5-b36f-b3e55cf620c2 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Asynchronous stochastic approximation and Q- learning,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59d2640a-5d45-44aa-b9ff-1e4605571daf · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Learning without state- estimation in partially observable Markovian decision processes,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f091b0cf-ee18-4f4a-bb48-fc5ff88244e9 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Gradient-based algorithms for zeroth- order optimization,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df21a8b9-1592-4b0e-a279-2b28286661cd · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d962d91-cbdd-4605-aec2-ea7843db0b2c · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 861e7c8a-51fa-45db-8e7c-2d55f55dcccb · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs future states
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5a98ed7-d2b8-4cf5-b4cb-df55b3d781f7 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f606a3a-2faf-468d-add3-72e98cb0c788 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9b99d8f-132e-4671-ad10-b449dc842323 · outbound
Convergence of regularized agent-state-based Q-learning in POMDPs These two cases must be considered separately
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.