Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2506.01639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01639 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.550668Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 037ab0cc-4bfe-4cb9-83c5-532dc0ba7cdc · outbound

This paper cites Deriving and improving cma-es with information geometric trust regions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Deriving and improving cma-es with information geometric trust regions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.449301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.434436Z digest=sha256:ae415af92de670a985ed29ee339886d8ab6fe1809368b21338d3e90d079f53d6

Observation e99c98b3-4569-4f1f-8ab5-30ed82433c2b · outbound

This paper cites Maximum a posteriori policy optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a posteriori policy optimisation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.438011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.454596Z digest=sha256:9f2a80293b8ce05afe18a3022d753e8bdb2eafcc09730bc435c9a36d752995de

Observation edc3fbf0-fb53-47d0-ad32-f868c9c1eef6 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.475700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.475700Z digest=sha256:2d7f47c1400454fb8e394ccc9544404b7eac28ec973ee9abb734edaa891123a2

Observation 669f855a-fb1f-4345-8f34-9131c79842e1 · outbound

This paper cites Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.427993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.497612Z digest=sha256:29abb141979b948b9e49bc6416c9909ee7bb49f159a9b6b3636919d66874399e

Observation 998ab124-4c61-4dbe-9271-1092e56717df · outbound

This paper cites Sukhatme.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sukhatme

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.418179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.519432Z digest=sha256:d15113227081dae58a3d34889f989a2ec3dede54b0ea4e4823ccf50a53f8d38d

Observation a76f75fc-75d7-4c36-86b2-3f2e8582a2bd · outbound

This paper cites Box2d: A 2d physics engine for games.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Box2d: A 2d physics engine for games

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.408208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.548555Z digest=sha256:87b69b18baf497a1c18aaa5e193f0e7c1643b41b59a98ec6ade9c86b12750d11

Observation 91e8e127-d4ff-421f-950a-dd767d09f0a5 · outbound

This paper cites Greedification operators for policy optimization: Investigating forward and reverse kl divergences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Greedification operators for policy optimization: Investigating forward and reverse kl divergences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.569756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.569756Z digest=sha256:68ddd76d945a13059f5840bb64292702c7a500a7508be85dc3a20c54d01402d3

Observation c3becc7a-769b-42c3-9abb-f51c1c418bb5 · outbound

This paper cites Using expectation-maximization for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Using expectation-maximization for reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.392907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.606210Z digest=sha256:aa3d7feac94750b1f6bc65fdc505f6c1a4536867ae9687c1011e7a79b28001fb

Observation a15b1794-9e53-4f58-a320-1a8848be737a · outbound

This paper cites Soft actor-critic for navigation of mobile robots.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic for navigation of mobile robots

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.383303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.642675Z digest=sha256:160f3c262f92c6a1a0143857bc973743d71c53d53b7a79767496b98816c93b32

Observation 4c72111c-d7f2-4dc9-991f-b248de50cd2f · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.374373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.671122Z digest=sha256:56fe27131f3f583b8bfb90ebf17b5281d0f0f08d9c1f96ec15ff2373f27218c3

Observation 9e1a854b-3c7a-4867-a536-e8a699640a70 · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.364809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.692111Z digest=sha256:e0aa976991509272bf83ec234ef2f5cbef8446383713a360e7d268fb490b6a6b

Observation a0d0cc48-fe5e-45cc-9910-719dcbb4657c · outbound

This paper cites Virel: A variational inference framework for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Virel: A variational inference framework for reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.355409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.713449Z digest=sha256:9c73cb513afcd777a8d272d1565b4c91ba0300615eee73493086c4824ee9da7c

Observation 3032c6c8-423b-4b27-b522-9467f96829ad · outbound

This paper cites Brax - a differentiable physics engine for large scale rigid body simulation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Brax - a differentiable physics engine for large scale rigid body simulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.345980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.734844Z digest=sha256:d67243dcfa5274ada90da9a7d9cc8d8d8a5fdd735b3d2eff9a661ff84013a298

Observation cfe213d7-5062-4740-bba8-ddca2d989ba3 · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Iq-learn: Inverse soft-q learning for imitation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.336235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.756905Z digest=sha256:ec812a1b44d99d38352535a689123998ed91e09e3d9185c55f2e7f0e3eb834ab

Observation 589bee85-2e5e-4e9a-9106-8d32a3887ca8 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Reinforcement learning with deep energy-based policies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.326371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.779818Z digest=sha256:508d0ef2fddfcdc7489cb9154350925f9b004fdb86a4820e6eae6e5339e6bb26

Observation 13f7922a-3653-4d92-8f01-029b477554f0 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.317218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.801739Z digest=sha256:b45dbdb6487f036e9444de201bfdf6643a3d91142b106f54cc94abcef5bbc63b

Observation 09af56c0-619f-4464-b928-8b977479ce74 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.823457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.823457Z digest=sha256:339c5eb66198d0fba907f78677ba1847e0f442f2635ca58f15dd978101801d46

Observation 07c4f6a2-5767-458f-b939-a5c2a23099d9 · outbound

This paper cites The curse of dimensionality for numerical integration of smooth functions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning The curse of dimensionality for numerical integration of smooth functions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.307790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.844042Z digest=sha256:22f3d8b96df14980500c02dce2a255a7813a6f270d0b5b067f6d46b707a47f2d

Observation 6f661f2b-2f62-430a-b6b9-17b221abea52 · outbound

This paper cites Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.297605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.874324Z digest=sha256:2ab476281930ee1ea68ac6b4a68965bfa6d2470bdec0d96bd0691669a72ad55a

Observation c1ada51c-5566-4cae-8b05-ee776a377e36 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.288646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.903554Z digest=sha256:1f54fa1692f57ef7a38a7dbd13fa5e967a5060a1ad7c411865de7c90d7fdbd42

Observation 9ee33d03-f8d9-4270-a3e6-bdcd64224198 · outbound

This paper cites Accelerating reinforcement learning with value-conditional state entropy exploration.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Accelerating reinforcement learning with value-conditional state entropy exploration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.278864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.931033Z digest=sha256:647d2da9fc5c6b95be8641a7d5155e13a7e7120cc0e9a915497c4ae37389fda3

Observation fff43b0d-d505-4b4c-bb31-b943f546a04b · outbound

This paper cites Optimistic reinforcement learning by forward kullback--leibler divergence optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Optimistic reinforcement learning by forward kullback--leibler divergence optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.268488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.952445Z digest=sha256:8a408d26b6fe1a460535502d245c7b5975bbeeb461717817732457f61ac2e118

Observation f4940d73-37ee-4291-af6e-d085bf9bb700 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.258046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.973438Z digest=sha256:161dc68c2be589f785a69ead5d39fccebd8ec5f90a0bba6e5ce778e73f4a2025

Observation af218bc2-7795-4bd9-99b4-11c5e994d2ce · outbound

This paper cites Continuous control with deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Continuous control with deep reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.995064Z digest=sha256:52fd3f8bd1457fc753112309cbfb13e6a5a4ea7cac8b9dcd901a25275b600577

Observation 7a4ee0e9-c39c-4524-871a-32c1d5b62299 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Constrained variational policy optimization for safe reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.248022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.017398Z digest=sha256:262a6066dba02937f332083a8927bab888fe712602052000970db9161d1e363b

Observation ef99ea56-bcce-4636-9b5a-a19d99e946b8 · outbound

This paper cites Algorithm 145: Adaptive numerical integration by simpson's rule.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Algorithm 145: Adaptive numerical integration by simpson's rule

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.042879Z digest=sha256:bd77d4856ba53a1520f98144dcc03c8b1c48c50d0315c90a1a284c5c59d84590

Observation 31cb562f-2668-4bdc-998a-82683a8e3736 · outbound

This paper cites On principled entropy exploration in policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning On principled entropy exploration in policy optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.227828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.064724Z digest=sha256:23400c213d4a9b399d83bfb79abdfac0614f7ca3f8e13db6200e671432981f7b

Observation a4ce579e-74e0-4780-b0b3-f6b3ebf5c88f · outbound

This paper cites Importance sampling techniques for policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Importance sampling techniques for policy optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.217571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.095121Z digest=sha256:8419ec01c63dfe9c9519f78bfe1f057734c2d129d6a35c25364037a96b3d4ca7

Observation ebb7a7e3-ef59-4277-9405-574668b10395 · outbound

This paper cites Human-level control through deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Human-level control through deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.147276Z digest=sha256:6673d493698e5ab0cbcd0af17f1f651cca7be2abd0dcf2563c55c8431b0c7fb8

Observation 9198cf18-ceb7-4813-9cd9-0f3804966b78 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.201661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.170202Z digest=sha256:f85602c3b2eb30fb448102e56c72f94555e61c8248ac788ab99bb464d8bd0fca

Observation c4b94ee6-79d0-4931-8ac2-ba475ccf4576 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.196122Z digest=sha256:7275d5be3a3c0d4ab840ec0cab30fe4548c5f24f1fb1671ac71422e0cfff4824

Observation c225e167-9421-4220-b4f7-27df15930022 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.217785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.217785Z digest=sha256:3e17f64353fe235b4b8a16e2601769511630d876e8f8e0c1da59923326109a16

Observation 69f38043-f28e-468a-a69e-d6efaccbcd41 · outbound

This paper cites Trust region policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Trust region policy optimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.182420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.238985Z digest=sha256:a6a82e9a87418f54afae8a7dfd2b78cee851df6521374e195061a994b522a707

Observation a83300ff-a461-4fae-ab21-f7d0736b5e1d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.261078Z digest=sha256:2a066653be78763ffe05680e64fc1fbfb1546ab06e14f23fe70924947da3cf04

Observation 1513d567-505b-47af-9015-2f7aca0c2c7d · outbound

This paper cites Monte carlo sampling methods.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monte carlo sampling methods

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.134269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.285541Z digest=sha256:0a211fe73c0353af49c1f123c0470cc1f673280cdd8d48a2d3dac7786ebc8c9b

Observation 10f56de9-2b04-4e8c-af7e-3c56428a2a02 · outbound

This paper cites V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.056253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.312110Z digest=sha256:5811dbec026ab7089d67ed8006767a60e4bdc044429d082033ec3cb8d40e885e

Observation 884cde22-32a5-463b-b479-c70cad6a28a7 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.333612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.333612Z digest=sha256:20a5030b319b8423b72ff6679e6c14668539380b304474aa9220df76e90815b1

Observation a0e0493e-d1ca-4cdc-96c2-5fbe9793ae0a · outbound

This paper cites Sutton and Andrew G.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sutton and Andrew G

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.371554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.371554Z digest=sha256:34a7cd59c1832b4a0ada6ca803e951685e1eaadee144b913c6bb679fd1c58c7f

Observation e1a785fa-c6f7-4a64-910c-c4b7f60424ba · outbound

This paper cites Mujoco: A physics engine for model-based control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.393604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.393604Z digest=sha256:eb2444731bedea2a1a4f50405064b8c775880322d94a3a8cd2eca97b2fc883a9

Observation 4872ae71-a7d7-4117-b3db-df8724e4ae0b · outbound

This paper cites Probabilistic inference for solving discrete and continuous state markov decision processes.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Probabilistic inference for solving discrete and continuous state markov decision processes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.998747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.414592Z digest=sha256:65f768f0c937b583a7cec944514e7b8c1dabe310ccf24795b2d4f950c5143854

Observation e2e48a29-57ef-40ea-b092-25ed2df07db4 · outbound

This paper cites Self-play reinforcement learning guides protein engineering.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Self-play reinforcement learning guides protein engineering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.939905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.436726Z digest=sha256:5a46340f81609705322cff8b3a29be60f80d4748db2ff16f01105494f6dc61fb

Observation 1b92f031-66b6-4596-a6af-efae17983045 · outbound

This paper cites Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.897254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.458763Z digest=sha256:e5645df6c026f7d1593e452eb72d649345a72d37b9b88953a5aeb416f33d8cff

Observation d012f224-7276-4e29-92e3-93359c316e22 · outbound

This paper cites Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.839762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.480133Z digest=sha256:230697fde30b3c33ff44d5e4c960a90083b22c23d80cab8b6b5e9e90703d877c

Observation 1bdd5f55-a73a-40ce-a820-4a8fafe4d91e · outbound

This paper cites Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.785426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.501065Z digest=sha256:b5ba9ab6aefbc9dac1d6b4bb33fc1eb1976c1e93db9e291e9cb457cc8dd64d5e

Observation 84d7064f-7693-4298-be96-8accd733b318 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.731567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.523235Z digest=sha256:a3cd1016d61d40794d159f419c6ca7a645dfc4fe3fdf6fa44f1699749be745d8

Observation 7cb4a608-aa9c-4c54-b024-3d59f93e49df · outbound

This paper cites Wasserstein gradient flows for optimizing gaussian mixture policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wasserstein gradient flows for optimizing gaussian mixture policies

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.688869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.550668Z digest=sha256:9bc340e0f2f56559eec0274554804133f41902901bff2c3d8621634f082492cc

Pith citing papers

No inbound Pith citation observations are available.