Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.550668Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2506.01639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.550668Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 037ab0cc-4bfe-4cb9-83c5-532dc0ba7cdc · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Deriving and improving cma-es with information geometric trust regions
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e99c98b3-4569-4f1f-8ab5-30ed82433c2b · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a posteriori policy optimisation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation edc3fbf0-fb53-47d0-ad32-f868c9c1eef6 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a Posteriori Policy Optimisation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669f855a-fb1f-4345-8f34-9131c79842e1 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 998ab124-4c61-4dbe-9271-1092e56717df · outbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a76f75fc-75d7-4c36-86b2-3f2e8582a2bd · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Box2d: A 2d physics engine for games
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91e8e127-d4ff-421f-950a-dd767d09f0a5 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Greedification operators for policy optimization: Investigating forward and reverse kl divergences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3becc7a-769b-42c3-9abb-f51c1c418bb5 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Using expectation-maximization for reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a15b1794-9e53-4f58-a320-1a8848be737a · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic for navigation of mobile robots
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c72111c-d7f2-4dc9-991f-b248de50cd2f · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e1a854b-3c7a-4867-a536-e8a699640a70 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0d0cc48-fe5e-45cc-9910-719dcbb4657c · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Virel: A variational inference framework for reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3032c6c8-423b-4b27-b522-9467f96829ad · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Brax - a differentiable physics engine for large scale rigid body simulation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfe213d7-5062-4740-bba8-ddca2d989ba3 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Iq-learn: Inverse soft-q learning for imitation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 589bee85-2e5e-4e9a-9106-8d32a3887ca8 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Reinforcement learning with deep energy-based policies
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13f7922a-3653-4d92-8f01-029b477554f0 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09af56c0-619f-4464-b928-8b977479ce74 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft Actor-Critic Algorithms and Applications
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c4f6a2-5767-458f-b939-a5c2a23099d9 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning The curse of dimensionality for numerical integration of smooth functions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f661f2b-2f62-430a-b6b9-17b221abea52 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c1ada51c-5566-4cae-8b05-ee776a377e36 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ee33d03-f8d9-4270-a3e6-bdcd64224198 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Accelerating reinforcement learning with value-conditional state entropy exploration
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fff43b0d-d505-4b4c-bb31-b943f546a04b · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Optimistic reinforcement learning by forward kullback--leibler divergence optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4940d73-37ee-4291-af6e-d085bf9bb700 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Conservative q-learning for offline reinforcement learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af218bc2-7795-4bd9-99b4-11c5e994d2ce · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Continuous control with deep reinforcement learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4ee0e9-c39c-4524-871a-32c1d5b62299 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Constrained variational policy optimization for safe reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef99ea56-bcce-4636-9b5a-a19d99e946b8 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Algorithm 145: Adaptive numerical integration by simpson's rule
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31cb562f-2668-4bdc-998a-82683a8e3736 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning On principled entropy exploration in policy optimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4ce579e-74e0-4780-b0b3-f6b3ebf5c88f · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Importance sampling techniques for policy optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebb7a7e3-ef59-4277-9405-574668b10395 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Human-level control through deep reinforcement learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9198cf18-ceb7-4813-9cd9-0f3804966b78 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Asynchronous methods for deep reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4b94ee6-79d0-4931-8ac2-ba475ccf4576 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improving Policy Gradient by Exploring Under-appreciated Rewards
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c225e167-9421-4220-b4f7-27df15930022 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f38043-f28e-468a-a69e-d6efaccbcd41 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Trust region policy optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a83300ff-a461-4fae-ab21-f7d0736b5e1d · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1513d567-505b-47af-9015-2f7aca0c2c7d · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monte carlo sampling methods
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10f56de9-2b04-4e8c-af7e-3c56428a2a02 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 884cde22-32a5-463b-b479-c70cad6a28a7 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e0493e-d1ca-4cdc-96c2-5fbe9793ae0a · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sutton and Andrew G
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a785fa-c6f7-4a64-910c-c4b7f60424ba · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Mujoco: A physics engine for model-based control
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4872ae71-a7d7-4117-b3db-df8724e4ae0b · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Probabilistic inference for solving discrete and continuous state markov decision processes
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2e48a29-57ef-40ea-b092-25ed2df07db4 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Self-play reinforcement learning guides protein engineering
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b92f031-66b6-4596-a6af-efae17983045 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d012f224-7276-4e29-92e3-93359c316e22 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bdd5f55-a73a-40ce-a820-4a8fafe4d91e · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84d7064f-7693-4298-be96-8accd733b318 · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum entropy inverse reinforcement learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7cb4a608-aa9c-4c54-b024-3d59f93e49df · outbound
Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wasserstein gradient flows for optimizing gaussian mixture policies
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.