Pith. sign in

Paper Citation Record · LEDGER

Soft Actor-Critic Algorithms and Applications

As of 5 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 100 inbound Pith citation observations for arXiv:1812.05905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1812.05905 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T14:32:41.948156Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 128 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:55:18.610364Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact10
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

1955
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 120c6f10-402a-451f-a694-cd7837bfd158 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Soft Actor-Critic Algorithms and Applications Maximum a Posteriori Policy Optimisation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:41.989748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:c807f0582382f8688c24d6dbfb40f93d40c5a1ad2b45a88f0002e096c2fd7ee7

Observation db963ab7-95cf-4745-a90f-af682c73a1b5 · outbound

This paper cites OpenAI Gym.

Soft Actor-Critic Algorithms and Applications OpenAI Gym

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:32:42.037081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:a2c02d9bfab35547470d6e630a822ab10f27f0f6cd4c8a87918a98bf447ff6fa

Observation 3c434479-d9c5-428f-8cd0-956793b96077 · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods.

Soft Actor-Critic Algorithms and Applications Addressing Function Approximation Error in Actor-Critic Methods

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.002405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:97ce3f8cdcb1f94ce248d8baad842f86c6d4b65e2f722f200a85e34ec6c887f8

Observation 6214e5e7-2d98-4b8c-a768-3ca46054c422 · outbound

This paper cites The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning.

Soft Actor-Critic Algorithms and Applications The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.013123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:a1a3225c8d64101b60d4ab892f97812cacb5a5a21f61d2f0fa6e8a89f53ead28

Observation 9a3daa3f-8c80-4580-aaf7-70243aa818dc · outbound

This paper cites Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic.

Soft Actor-Critic Algorithms and Applications Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.020468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:89c5e78c2cf0cc46e0235496ae7770de8591987c97c18e5d93da4f40b87bee3a

Observation 4d1bfb09-22c5-4c4d-9d21-488274632bc0 · outbound

This paper cites Deep Reinforcement Learning that Matters.

Soft Actor-Critic Algorithms and Applications Deep Reinforcement Learning that Matters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.027745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:5e8816548d7bc3690a5bb70368701ad251de2f2986fd81523adcea60a2568e67

Observation c649c500-bee0-4e4b-b875-47f8f39a874f · outbound

This paper cites an unresolved cited work.

Soft Actor-Critic Algorithms and Applications Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-13T14:32:42.076074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:43929f86b2b04195245aae6f24932ff3a2e03efbb5c2dd68e38a7618ab1d750c

Observation 6b3b9993-beef-4bd2-8bc5-0ae548031259 · outbound

This paper cites Continuous control with deep reinforcement learning.

Soft Actor-Critic Algorithms and Applications Continuous control with deep reinforcement learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:32:42.043067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:cb3625387e6aaeb59eccf63e4e763235d9abda9ead963f5574068cb3067d8256

Observation a60cd17a-8fb2-439a-a9cb-89331ef8be1b · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Soft Actor-Critic Algorithms and Applications Playing Atari with Deep Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:32:42.050839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:4d38d71798bf9afce8bfc45429fc14533e35a0d48f905235cd810e2ac6e13c8b

Observation e9513ef5-76ce-46d4-aa1f-fe42ab84d8da · outbound

This paper cites Trust-PCL: An Off-Policy Trust Region Method for Continuous Control.

Soft Actor-Critic Algorithms and Applications Trust-PCL: An Off-Policy Trust Region Method for Continuous Control

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T22:34:15.110658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:b8632ee167bf8c5a0a1f79fcc9bb0922f7a9a849538746c7685284ca71d82a43

Observation ed83ccbc-3954-48cc-bde3-f1919af2aca1 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

Soft Actor-Critic Algorithms and Applications Equivalence Between Policy Gradients and Soft Q-Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.065863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:282bfeade74d8cd2cef78db49d14c860df807f67a5847ca46201d23db907e97c

Observation c4c0db9e-1043-4521-a4a0-8772abbadb48 · outbound

This paper cites Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost.

Soft Actor-Critic Algorithms and Applications Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:41.977302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:6c771916b043ca3928a1b9e8ed1d1d6c09bf886bd734f2bb35f163343cc723f9

Observation 37a23961-452a-4e81-8400-e5b869491c7f · outbound

This paper cites In that sense, discounted policy gradients typically do not optimize the true discounted objective.

Soft Actor-Critic Algorithms and Applications In that sense, discounted policy gradients typically do not optimize the true discounted objective

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T14:32:42.070812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:32:41.948156Z digest=sha256:7fe95b584b54f469e6d99570e040b4cf85040f821a041b9236ae1cc5e6b8600c

Pith citing papers

Observation 7da5b571-da5e-456f-877b-c72e364a954e · inbound

Solving Rubik's Cube with a Robot Hand cites this paper.

Solving Rubik's Cube with a Robot Hand Soft Actor-Critic Algorithms and Applications

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:38:28.936381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:38:28.621842Z digest=sha256:8b955ad1503574092585085564317622367b1c90b8d0fb25a211eb02f89612a4

Observation 0b5d07ec-d09a-4f32-8783-5397d40a35de · inbound

robosuite: A Modular Simulation Framework and Benchmark for Robot Learning cites this paper.

robosuite: A Modular Simulation Framework and Benchmark for Robot Learning Soft Actor-Critic Algorithms and Applications

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T22:37:13.401815Z digest=sha256:a8cb6c047ae84e7918150f436007508c384a0cc83d78b5ae37f8c21422367499

Observation 1453e9c2-b39d-4c48-a384-d93b0f74a5a3 · inbound

FP-IRL: Fokker--Planck Inverse Reinforcement Learning -- A Physics-Constrained Approach to Markov Decision Processes cites this paper.

FP-IRL: Fokker--Planck Inverse Reinforcement Learning -- A Physics-Constrained Approach to Markov Decision Processes Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:29:11.424793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T08:26:33.267655Z digest=sha256:ff3d5c2b8c7bc166956aba33d9df3a6cc748f02a44d057a10ec626dc018a9f99

Observation 0387d6f7-4683-4d6b-abad-47135f081e56 · inbound

TD-MPC2: Scalable, Robust World Models for Continuous Control cites this paper.

TD-MPC2: Scalable, Robust World Models for Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T17:27:35.932371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T17:27:35.733800Z digest=sha256:0f8c51660e6ad0692d5dd2a50d0167b83961b8519235f859f592217ffbe5aa3a

Observation e2cbd209-c7fb-4a05-b9ce-5170d3ce790d · inbound

CROP: Conservative Reward for Model-based Offline Policy Optimization cites this paper.

CROP: Conservative Reward for Model-based Offline Policy Optimization Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:39:01.278088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-24T06:37:57.104233Z digest=sha256:b948c9a334d4c5aedec022f16167eedcd1d368fb3f569b0cae4a538f4211d773

Observation 0b010a92-4cb6-4395-ae42-61d86fdc3713 · inbound

Koopman-Assisted Reinforcement Learning cites this paper.

Koopman-Assisted Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:53:48.798199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T02:49:43.717890Z digest=sha256:28c63efc9c3b373f740bb7787106d6a461dccb8f66296a17bdee0502485d43e3

Observation e09dadb6-cde6-4d75-9b6d-d041c992053f · inbound

Optimal Gait Control for a Tendon-driven Soft Quadruped Robot by Model-based Reinforcement Learning cites this paper.

Optimal Gait Control for a Tendon-driven Soft Quadruped Robot by Model-based Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-23T23:53:38.718102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:51:21.019535Z digest=sha256:ff3fa007fce15f25eb633ce3f9c3b4b82a77d223f9964b7d2041f0a3d4d4121b

Observation c2b80a4c-f453-45f2-9592-6320e793fbbe · inbound

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting cites this paper.

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:52:26.623292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T02:49:40.277048Z digest=sha256:0e70c4f135fd9e96443f34f511ea7b63295700b3c77c62abff204ac44cfaf56f

Observation b9a4d2fc-82d9-45f8-b5ea-511557d7374a · inbound

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning cites this paper.

Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:51:45.699715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:49:44.263123Z digest=sha256:cdcd705dc2246f4f9fbb26ad1a0e6362f0df03810504bb3f3dc31ea2169d357a

Observation 9760c728-ec29-4196-9e4b-5c7a7e81932b · inbound

Accelerated Learning with Linear Temporal Logic using Differentiable Simulation cites this paper.

Accelerated Learning with Linear Temporal Logic using Differentiable Simulation Soft Actor-Critic Algorithms and Applications

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:52:15.118525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T10:50:18.047151Z digest=sha256:58f1c75cc6d0a09259396ff94ebb94baf5c018434f8ec4e06c5ecc99893cb11b

Observation e8e5e04d-8eb4-4849-94b5-d83b1ea8b6be · inbound

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty cites this paper.

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty Soft Actor-Critic Algorithms and Applications

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.675283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:09:51.531488Z digest=sha256:3822f8120c444d08c95ea9437990f9aead78de50a9b87a520a9545c289e18ba9

Observation 514b3a63-17a3-4436-9091-30655c9af17b · inbound

Deep Double Q-learning cites this paper.

Deep Double Q-learning Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:24:27.690993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T00:23:41.373753Z digest=sha256:7a9c1cb6b1f6a39fd4f1d79a1ae40d13ebfcc988c899d198a0d1bace6a9c8884

Observation 441c3157-77e6-4b7d-8c39-ea13c868c73b · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Soft Actor-Critic Algorithms and Applications

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:05.210606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:a0b30507b65e7fa96dcb56f76d6205b55cb06b0ed95a727c896849b371fd2d60

Observation 1ecd2410-400d-4ea5-9193-50e0c50907ba · inbound

Joint Scheduling of Deferrable and Nondeferrable Demand with Colocated Stochastic Supply cites this paper.

Joint Scheduling of Deferrable and Nondeferrable Demand with Colocated Stochastic Supply Soft Actor-Critic Algorithms and Applications

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:52:03.701273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T04:49:54.452877Z digest=sha256:de140623bce98001938b1bb5d7d952758d4ad63dc144f3f6408305bd8914c03e

Observation 0926d32f-33a8-4238-a6a1-9a2b6d55eefe · inbound

Relative Entropy Pathwise Policy Optimization cites this paper.

Relative Entropy Pathwise Policy Optimization Soft Actor-Critic Algorithms and Applications

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:22:04.030185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T04:19:15.018380Z digest=sha256:7a79edfdcf5d58c4d0333867d9fef2c7a0c16dc864c6a847f1f68ba2f95cc03f

Observation 3c12a234-e4d3-4de2-b97d-61f99e7fb5d7 · inbound

Adaptive Ensemble Aggregation for Actor-Critics cites this paper.

Adaptive Ensemble Aggregation for Actor-Critics Soft Actor-Critic Algorithms and Applications

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T02:16:59.354135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T02:14:26.752215Z digest=sha256:d061e5e17104bd9c0b3780cc6d142273feb4bc59cbdc183e6f7309b7a854d867

Observation 532e41a7-e47f-452f-afcb-f6236261a1d2 · inbound

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies cites this paper.

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies Soft Actor-Critic Algorithms and Applications

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:06:57.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T01:05:17.086399Z digest=sha256:e542296f05ff83b06aa9b1f6008b4f4b61619ba72db29e6b78b2aabfa7290589

Observation dacb7721-7cc3-412a-aec8-d89b4dd0d079 · inbound

First Order Model-Based RL through Decoupled Backpropagation cites this paper.

First Order Model-Based RL through Decoupled Backpropagation Soft Actor-Critic Algorithms and Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:55:18.610364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:55:18.610364Z digest=sha256:f4e7706a956b216a9e074ae60025822ef0cec455cb01330512f16895be9518ec

Observation f4ab26b5-cdc7-4628-ab95-277411eca475 · inbound

Pinching Antenna System (PASS) Enhanced Covert Communications: Against Warden via Sensing cites this paper.

Pinching Antenna System (PASS) Enhanced Covert Communications: Against Warden via Sensing Soft Actor-Critic Algorithms and Applications

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T00:07:52.792545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:07:52.792545Z digest=sha256:fe281b0e8358b7ac5e391630dd31415c3906e20eb74b6185a0d11f3f5599a5b6

Observation c81b2da2-632d-453e-b65a-b60dc87ec669 · inbound

Attention and Risk-Aware Decision Framework for Safe Autonomous Driving cites this paper.

Attention and Risk-Aware Decision Framework for Safe Autonomous Driving Soft Actor-Critic Algorithms and Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:03.122664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:03.122664Z digest=sha256:7e91e11c9da04041143a20a306d6ade453c217e41fa679576c15b6914b9925b0

Observation ecb03fdc-7934-4344-b447-05ecf8c4d8c3 · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Soft Actor-Critic Algorithms and Applications

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:06:39.748856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:d4fa3d4229175c81f6f118fe4063a36607e806e3c44c4e214cfbe242736db858

Observation cb6c0065-2f26-432c-a2fa-7dd6b8c9e369 · inbound

Real-time reinforcement learning for turbulent state-dependent control in a bluff-body wake cites this paper.

Real-time reinforcement learning for turbulent state-dependent control in a bluff-body wake Soft Actor-Critic Algorithms and Applications

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T22:44:24.425922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T22:42:00.056915Z digest=sha256:f2697c69bc46cdf8c2b1e19135f64c426864d12155232942624baf53864b18db

Observation 05af16a5-ca92-4f9a-8214-d4dd341bcb93 · inbound

CORB-Planner: Corridor as Observations for RL Planning in High-Speed Flight cites this paper.

CORB-Planner: Corridor as Observations for RL Planning in High-Speed Flight Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T16:56:47.457432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:56:47.457432Z digest=sha256:81076c6fb83679f3c869ac2d0f37ecf3a930c34055ec7ad93a78955aee7a2e2a

Observation c6996be6-81ad-4798-8b84-3b0317f39487 · inbound

High-Precision and High-Efficiency Trajectory Tracking for Excavators Based on Closed-Loop Dynamics cites this paper.

High-Precision and High-Efficiency Trajectory Tracking for Excavators Based on Closed-Loop Dynamics Soft Actor-Critic Algorithms and Applications

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:31:33.450845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T15:30:19.253933Z digest=sha256:02260a30417721977f05abb5dc53090af9dbcbe41f93b3fd64c36bf7d2e26b5e

Observation bfb88eb3-2014-44bd-b51e-3d4650a46604 · inbound

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning cites this paper.

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:16:21.599517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:15:50.805708Z digest=sha256:a195bd8b097a7af5cbe423f7d3b4de7185992e2ad3fc045ef34ca0c94533a973

Observation 358cb053-1672-455c-8516-d252540d03e6 · inbound

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning cites this paper.

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T12:16:21.481489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T12:15:50.805708Z digest=sha256:655365001489b0340ba15b08c04a37372c35468a0562072caccc641746e71a07

Observation f473e237-c902-4191-878e-2294cefac0a7 · inbound

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning cites this paper.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.323552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:cec91542fe81f67e3c3d649fd0f87e48b858a8e456ca7ec4d4910f1dfe9cf3a0

Observation 9e3c1d92-83e3-4ca5-88ba-e1ecd28a8324 · inbound

Optimal control of the future via prospective learning with control cites this paper.

Optimal control of the future via prospective learning with control Soft Actor-Critic Algorithms and Applications

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:05:25.321470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T23:04:54.956095Z digest=sha256:58a5d38fe8114f42ac7ac0bb05661f437b2eb05d59e42028ed0d34f48350946e

Observation 73d398b9-9885-420e-993b-52bf937ebf85 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:05:15.565069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:04:41.766964Z digest=sha256:1d3800ea5cf0368f70c1435c32acaac86765dcbcb253a44bf8f8ddcd363211cd

Observation f6c6dc51-1296-4c12-98e6-e1a0679284c3 · inbound

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cites this paper.

Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T21:40:43.870053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:40:43.870053Z digest=sha256:68f124f50ef98271af6e8d69f83f1e13ffda8e309a20db79b084c913250ddcd9

Observation fcdc9dcd-3e93-43ce-bb7e-19946af7fa7c · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:50:15.261334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:45:15.034196Z digest=sha256:0c756fd02f1164488b07edf7f143c909f0d996f94db3549d52c3aef454a67579

Observation a1643d85-eec6-468e-87a1-d17242d0dd14 · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:20:28.851507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:17:26.381232Z digest=sha256:913c412349e2a84010ece11c3d66cccaee480ecb085dcd135e97e9f1b87b2cf5

Observation 95d99f5d-bda2-4a7c-b026-08f969a348f9 · inbound

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability cites this paper.

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:20:11.604219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:19:24.841090Z digest=sha256:cf229a51676111256d2ee31425238f437b24fe2d3359fc29ebd4d9a06b3ad727

Observation b69198c5-b6ba-4006-961e-e3c6265b7ae8 · inbound

Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems cites this paper.

Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems Soft Actor-Critic Algorithms and Applications

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T00:38:44.983292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:36:26.202547Z digest=sha256:9e4fcd60611c7e650a1b5561a468a52c38ec4eda4303dff75802c3b0db78a1d8

Observation 9e980c30-3165-4cd3-bca0-3ba4620548e6 · inbound

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes cites this paper.

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes Soft Actor-Critic Algorithms and Applications

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:01:04.877390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T16:00:10.425309Z digest=sha256:90731f2f3fbc63d304230ef2ee87bd2c3dd49930df09516249adf54f281c96fe

Observation f68c516d-3a0b-4499-8585-60b1fd0c7a06 · inbound

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes cites this paper.

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes Soft Actor-Critic Algorithms and Applications

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T11:45:47.594375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:45:47.594375Z digest=sha256:98ba35cab25b8cf680dc9d2f3a4e896865594c0c89d1b74d2811ab5be7e5c3be

Observation 74b6f97a-63fa-4c28-9916-1c04bc9487b0 · inbound

Coupling Smoothed Particle Hydrodynamics with Multi-Agent Deep Reinforcement Learning for Cooperative Control of Point Absorbers cites this paper.

Coupling Smoothed Particle Hydrodynamics with Multi-Agent Deep Reinforcement Learning for Cooperative Control of Point Absorbers Soft Actor-Critic Algorithms and Applications

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T11:28:07.479935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:28:07.479935Z digest=sha256:0c0a854032ed5b5956ef03624e2bc1c41e6dcfe6966149ae5298f01c1891ef06

Observation 17303f39-74c6-4d13-8187-ce9fbfcc348b · inbound

From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner's Tutorial cites this paper.

From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner's Tutorial Soft Actor-Critic Algorithms and Applications

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T10:52:32.302308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:52:32.302308Z digest=sha256:cc4c0753f25d8632c3dd6f2176ea39e03e6ca0b08c5c739a397752c4073eb119

Observation eefce765-f798-4859-8b44-f3855decb498 · inbound

Error Amplification Limits ANN-to-SNN Conversion in Continuous Control cites this paper.

Error Amplification Limits ANN-to-SNN Conversion in Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T06:57:05.568437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:57:05.568437Z digest=sha256:32a104cd99203f499fcb8be2762ceb6a6cbb4d1d71baeef5d981c8b2271ca2f5

Observation b401733e-c3c8-42a7-8dbd-bdf11baa7076 · inbound

Agile Reinforcement Learning through Separable Neural Architecture and Applications cites this paper.

Agile Reinforcement Learning through Separable Neural Architecture and Applications Soft Actor-Critic Algorithms and Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:13:00.293028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:13:00.293028Z digest=sha256:7d6ea5bae213bd746c4e08524c2c158af0746a7064ff59bcacba6bbae87491cc

Observation 246dd066-841b-41c7-a065-a2bd62a92113 · inbound

LC-SAC: Lyapunov-Constrained Soft Actor-Critic via Koopman Operator Theory for Trajectory Tracking and Stabilization cites this paper.

LC-SAC: Lyapunov-Constrained Soft Actor-Critic via Koopman Operator Theory for Trajectory Tracking and Stabilization Soft Actor-Critic Algorithms and Applications

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:47:40.482295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:47:40.482295Z digest=sha256:22a332be3e99dea7c82c1a478ff6d0f640767f9bf0bef3e2bf9a68b5a6e71334

Observation f273cae3-e53d-4d1c-8099-b10c5b7bf316 · inbound

Coupled Local and Global World Models for Efficient First Order RL cites this paper.

Coupled Local and Global World Models for Efficient First Order RL Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:03:56.421007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:03:56.421007Z digest=sha256:39804310be8484ebadd7904b864619dfbb62cb8a1d6ac9c1a98ef6c9edbc7253

Observation 7789059b-9a07-42cb-a407-bdfedfc5dea1 · inbound

Robust SAC-Enabled UAV-RIS Assisted Secure MISO Systems With Untrusted EH Receivers cites this paper.

Robust SAC-Enabled UAV-RIS Assisted Secure MISO Systems With Untrusted EH Receivers Soft Actor-Critic Algorithms and Applications

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:54:10.623520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:51:58.851812Z digest=sha256:058bee886412b4df2609d2c13198f37690bd1cb38b6b06b454510d6040e253e5

Observation 612cc045-a79b-4a03-ad1e-82dd76e199ed · inbound

Maximin Robust Bayesian Experimental Design cites this paper.

Maximin Robust Bayesian Experimental Design Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:15:30.097235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:13:06.058747Z digest=sha256:29065432b10a2370500fc6383ec2a7dcfb65359fd4bb5faa3a1a6bc4c0a5bf65

Observation 486c2dc2-3e05-461c-86b7-7b244acf5b49 · inbound

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation cites this paper.

Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:49:54.458492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:49:03.333757Z digest=sha256:245538fedd08e05228c090ece9c73f3fa84e4621c21a57a6acc494fc22d2f079

Observation 95851a73-591a-427b-802f-61c734bfc990 · inbound

DRL-Based Spectrum Sharing for RIS-Aided Local High-Quality Wireless Networks cites this paper.

DRL-Based Spectrum Sharing for RIS-Aided Local High-Quality Wireless Networks Soft Actor-Critic Algorithms and Applications

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T17:28:25.739838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:28:25.739838Z digest=sha256:81a1d0fa42f8bf00e23360331074bb92652f06c08391eca4bbfa69441744aa4d

Observation 7f36c4ba-1005-4714-82f6-deffa8506f90 · inbound

Delayed homomorphic reinforcement learning for environments with delayed feedback cites this paper.

Delayed homomorphic reinforcement learning for environments with delayed feedback Soft Actor-Critic Algorithms and Applications

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:48:08.241230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T18:44:26.633008Z digest=sha256:df00863086fe95702237c626b227d61c0dbf481fa3607748d1bd09a8329f62de

Observation d4321f1b-bd07-41da-ab1b-6dc4d73471b0 · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Soft Actor-Critic Algorithms and Applications

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:48:36.372298Z digest=sha256:b289afea5fb1d76a4a9a7d0e614b3d806e8917d062bc2d3179566b86d9a9fb1e

Observation e3da9749-3322-4bf0-b126-e5faa8493991 · inbound

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation cites this paper.

Incremental Residual Reinforcement Learning Toward Real-World Learning for Social Navigation Soft Actor-Critic Algorithms and Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T00:08:53.253285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:08:53.253285Z digest=sha256:04315e03af20e225b896370159b42cb451361f113e65b18cc985364f0540ffd6

Observation ad244b88-38e9-479f-a4f9-8c82dd39c708 · inbound

SafeMind: A Risk-Aware Differentiable Control Framework for Adaptive and Safe Quadruped Locomotion cites this paper.

SafeMind: A Risk-Aware Differentiable Control Framework for Adaptive and Safe Quadruped Locomotion Soft Actor-Critic Algorithms and Applications

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:00:42.108837Z digest=sha256:98765914b593c1d7c14f7edd216584a3b3593195b8413e793200564f37135a16

Observation 161cb248-9b4a-4ba7-a6b3-3289aa1add26 · inbound

New Scheme Adaption Strategy for Hyperbolic Conservation Laws cites this paper.

New Scheme Adaption Strategy for Hyperbolic Conservation Laws Soft Actor-Critic Algorithms and Applications

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T23:09:48.353960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:09:48.353960Z digest=sha256:e69399cb64197f0cf5c6877a7b237d658a8434ca2deded7e140ce4038abd2363

Observation 6a25bd6b-59ec-4436-a3dc-1a4cc8762a3f · inbound

Physics-Informed Reinforcement Learning of Spatial Density Velocity Potentials for Map-Free Racing cites this paper.

Physics-Informed Reinforcement Learning of Spatial Density Velocity Potentials for Map-Free Racing Soft Actor-Critic Algorithms and Applications

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:47:38.002725Z digest=sha256:a46630966673158538c9021b9a36a7cb60baf9948e1d58c0d0679e16ceacb387

Observation 7e004d75-46f9-47e9-9147-4cd2a2c46b10 · inbound

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation cites this paper.

Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation Soft Actor-Critic Algorithms and Applications

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:27:25.150807Z digest=sha256:37be1ba2cca1599ebf4d0f8f591b2b4cfcc8d7be5d9b677a74226bace22c2efb

Observation bcd90826-73e6-42d1-a4e3-db7bc3fc3c28 · inbound

Autonomous Diffractometry Enabled by Visual Reinforcement Learning cites this paper.

Autonomous Diffractometry Enabled by Visual Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:03:35.110730Z digest=sha256:6dec41576ad38081310ec592b3d8db8f15d98675bd6b8814ab797378d3ca9983

Observation 246ae235-502a-4748-8ab8-5f186da7c2d8 · inbound

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling cites this paper.

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T09:39:29.875977Z digest=sha256:98e8d2e48adcfc77f7d39dd4073e3ccb041851b6ce8ee7d92d6594b0c9250f98

Observation f429daf2-bae0-43f2-8041-b6cc0bf7ff14 · inbound

Self-Predictive Representation for Autonomous UAV Object-Goal Navigation cites this paper.

Self-Predictive Representation for Autonomous UAV Object-Goal Navigation Soft Actor-Critic Algorithms and Applications

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:31:49.766698Z digest=sha256:7238e679d0c64fa4e115b3c58e32dd4cf462e3a3789ed3134f1d5e5492dfe0d6

Observation 08555fc4-88c0-4ca8-baca-aee0bf2a8c7d · inbound

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning cites this paper.

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning Soft Actor-Critic Algorithms and Applications

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T15:39:17.505374Z digest=sha256:a07890ee6d9bf366bef3e94cf4d6a15878e0a2b28066a690a54bb187e945bee7

Observation 0dfdda28-ccf0-44f6-b255-fcc7cc77e079 · inbound

Atomic-Probe Governance for Skill Updates in Compositional Robot Policies cites this paper.

Atomic-Probe Governance for Skill Updates in Compositional Robot Policies Soft Actor-Critic Algorithms and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T11:12:30.501043Z digest=sha256:683e99134a8bdc47b2f8c80ba9390f8fddd423d25c67c6a914d8250780e85015

Observation a8d20b2b-5a54-45d8-b892-a8e20b0a844b · inbound

Atomic-Probe Governance for Skill Updates in Compositional Robot Policies cites this paper.

Atomic-Probe Governance for Skill Updates in Compositional Robot Policies Soft Actor-Critic Algorithms and Applications

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:25:26.291747Z digest=sha256:20caf5abe0b2cc1ca264046f7c56fde1eee903515a0446cc16439b2b0b82c4fd

Observation 50b569e9-fbc3-4852-9b7e-873121f41dbf · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:c3cdde282635571fb099bc15066f090f8bd69264037bd85b14f39a6dfee5079a

Observation c64ff3a5-6cab-4554-a6d5-fdbca793c2f5 · inbound

Counter-Dyna: Data-Efficient RL-Based HVAC Control using Counterfactual Building Models cites this paper.

Counter-Dyna: Data-Efficient RL-Based HVAC Control using Counterfactual Building Models Soft Actor-Critic Algorithms and Applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T16:38:22.468101Z digest=sha256:4f2f80a142c5739da4fb6eed69eef60cafaf259f50c8c421a04e6fa415b09493

Observation 3edbc64a-2ec4-4ecc-ac04-2cc95416e8ad · inbound

Offline Reinforcement Learning for Rotation Profile Control in Tokamaks cites this paper.

Offline Reinforcement Learning for Rotation Profile Control in Tokamaks Soft Actor-Critic Algorithms and Applications

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T14:49:07.819928Z digest=sha256:80283498fa00dbda7737ba61453a1252e765ec54279d6521732b0e122e39bc22

Observation 46b446e5-9acb-40bb-bceb-fb78c2e2e1cf · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:cd1ede8452444a9bc677f1c6e26ec5a60712996b907894bfa03beb66c110211b

Observation 7f6539ee-7775-4f35-8318-14e0b1b27013 · inbound

REAP: Reinforcement-Learning End-to-End Autonomous Parking with Gaussian Splatting Simulator for Real2Sim2Real Transfer cites this paper.

REAP: Reinforcement-Learning End-to-End Autonomous Parking with Gaussian Splatting Simulator for Real2Sim2Real Transfer Soft Actor-Critic Algorithms and Applications

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:30:06.120140Z digest=sha256:8994bac76153435400e4ab46642a73a6bcad6642eeb9ad705e6aac3bf65fcaba

Observation f728f5d6-0d94-4980-8799-36e8b2609b69 · inbound

Generative Actor-Critic with Soft Bridge Policies cites this paper.

Generative Actor-Critic with Soft Bridge Policies Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:01:08.896356Z digest=sha256:a16744554d6a866b23b976518c714ae60082fe367a764bdcb6ef3c599792e387

Observation e11322e0-4c08-4519-9884-45916e93b863 · inbound

A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets cites this paper.

A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets Soft Actor-Critic Algorithms and Applications

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:21:30.413575Z digest=sha256:8b76853ae9f6ddd9824a84a254803bd7c78c80b15cbf562533970eb54a6f0120

Observation ab3fee38-09ea-4d14-8349-e644b5aee0e6 · inbound

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic cites this paper.

Revisiting Mixture Policies in Entropy-Regularized Actor-Critic Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T04:33:08.990277Z digest=sha256:dc2734acced21d5fcea8ff04a32f946ac281b3da94064dfb04e0305f351dff7e

Observation 7e07ff32-fd7a-4938-90f1-87ac4ebf505a · inbound

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency cites this paper.

TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency Soft Actor-Critic Algorithms and Applications

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:50:20.040271Z digest=sha256:fb139ef86596e64343b8b7be4ec429b447d74bdf391ba369ed5045829ad5c3a3

Observation 90e442bd-16f2-4b8b-8e18-88ce7c9cc47c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Soft Actor-Critic Algorithms and Applications

Reference 100

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:c82c77f5acbafa593a6e3c673a1313a8b533b77c7564fbf5c3f5f0762c6ada91

Observation c059de7c-b138-4742-8226-202d6bbb6492 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Soft Actor-Critic Algorithms and Applications

Reference 99

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:19:55.810377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:55.810377Z digest=sha256:371022f2843c52a079ffd316cb998e401d95c0510cb78cdef68357849957c873

Observation 8dbbcbc8-35ec-49d4-ad5d-bdbd6a21e90f · inbound

Debiased Model-based Representations for Sample-efficient Continuous Control cites this paper.

Debiased Model-based Representations for Sample-efficient Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:32:42.077468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:09:01.132424Z digest=sha256:007d23ba48c4989ffd28e6197d0fce38dfd6eb2d82ea2b58e919a2e0613ef43b

Observation 7dbaf19c-c29e-41af-a8bc-c7fbe3ff6ca9 · inbound

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation cites this paper.

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:12:55.209929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:12:07.920183Z digest=sha256:86bc33e332b13bc64c1b522457ccea8a73e2e362f006d5545b8afc78ffd86019

Observation 60af9d34-8a1a-4a3a-b563-8b928eb23b68 · inbound

R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning cites this paper.

R2R2: Robust Representation for Intensive Experience Reuse via Redundancy Reduction in Self-Predictive Learning Soft Actor-Critic Algorithms and Applications

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:39:47.541673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:39:23.297684Z digest=sha256:07d94a84c51cb9cd97ede562abb6c45997f7e6543656839a7945e2d215cd78ac

Observation e5f0bbdb-9bdf-46a2-825c-660402716ff9 · inbound

Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients cites this paper.

Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:13:30.162045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T02:12:58.603012Z digest=sha256:6ce473962b4fa175a56548750519fa2f4775e40c8d8dd8b60648968c55a6465b

Observation 9c5eccaa-60df-4a4e-a737-9e2a990171ff · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Soft Actor-Critic Algorithms and Applications

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T03:08:59.492498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:db7400ecada7072b7fa0fc7149272597f6b92223e154d02ca6f330fe02d4251b

Observation be7ff438-f973-4d72-b57e-75db7d7bf3c0 · inbound

EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control cites this paper.

EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T19:03:39.761378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T19:00:09.451593Z digest=sha256:a3d435644a79e49085d61cdb6186fd435cc33cf9f406644a81cac2d49141091c

Observation f764aca8-ad6f-40ba-9bc5-752ad0dff51b · inbound

COOPO: Cyclic Offline-Online Policy Optimization Algorithm cites this paper.

COOPO: Cyclic Offline-Online Policy Optimization Algorithm Soft Actor-Critic Algorithms and Applications

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:13:18.043222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T13:11:16.568415Z digest=sha256:e9d13c3eb805df15b241d1b2aa4b9ce2b2693737031b19825df98b304ed60aa5

Observation 68510778-92a3-4b7a-95d0-7640fa2cc6e5 · inbound

Unleashing the Power of Tree-of-Thoughts for Edge-Enabled AIGC Service Provisioning cites this paper.

Unleashing the Power of Tree-of-Thoughts for Edge-Enabled AIGC Service Provisioning Soft Actor-Critic Algorithms and Applications

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:08:06.979925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T07:07:44.112674Z digest=sha256:8a60308ac620d74b03a7ec928db7c9b5410ef3b38c699a73d54cd8a3ad108772

Observation c1d18125-8e2e-43bd-8c0e-ead9a19d46fc · inbound

Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight Aggregation cites this paper.

Partial Fusion of Neural Networks: Efficient Tradeoffs Between Ensembles and Weight Aggregation Soft Actor-Critic Algorithms and Applications

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T07:01:11.761823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-22T06:59:09.799156Z digest=sha256:a9024f2bce295112b83024e4a9d7cceb23e7a8566330d5d2fe41948896ff2e71

Observation b0840afa-5357-4aab-939a-6edec6a4c150 · inbound

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control cites this paper.

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:00:22.305127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T04:57:20.268613Z digest=sha256:24ceb8da1a551bc526f4bf4ec4c97f85a976e9e281d0910117bebf7a229f3e2b

Observation eacdfee5-c726-48e8-9c5e-b414eb30ab37 · inbound

Goal-Conditioned Agents that Learn Everything All at Once cites this paper.

Goal-Conditioned Agents that Learn Everything All at Once Soft Actor-Critic Algorithms and Applications

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:00:21.526534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-25T04:59:48.867927Z digest=sha256:674f6b8ebccca303d974a31793c0cbbce361583690dffbf863e8584fffcbbd80

Observation a7e7025e-882f-41bf-8b02-e1785d540b8e · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:50:20.678853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:60abb7e0ecbc20d98145ffaeff3f6a9908c47908878deab5ed8c09add069ed7e

Observation 346445e5-14e9-4adf-812b-2e032678a10c · inbound

Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language cites this paper.

Word Class Representations Spontaneously Emerge from Successor Representations Trained on Natural Language Soft Actor-Critic Algorithms and Applications

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:54:44.099998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T13:46:16.645506Z digest=sha256:ffd88a5cba8b4f14ed2b070a58a067be251c3310b7465d5ec5dc9451e4940278

Observation a6b07761-01a5-4046-a14f-e24972cf40f7 · inbound

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion cites this paper.

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion Soft Actor-Critic Algorithms and Applications

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:05:49.704942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:54:12.099045Z digest=sha256:ef800f836188ad56143dae642fcb46c9e609cd8d6b3c90ef225f2784c4f95f81

Observation a3802ecb-525b-4e2a-b5d7-ad2cb187694e · inbound

ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation cites this paper.

ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation Soft Actor-Critic Algorithms and Applications

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:35:50.717714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:26:22.777878Z digest=sha256:5df55b6ad06c0d3ea48a34fd92ed3a8ffa58592cd29e93b529855915ea62a934

Observation d809ab9a-ecb5-4227-8481-a457ed53e326 · inbound

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models cites this paper.

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models Soft Actor-Critic Algorithms and Applications

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.948067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:10:08.682307Z digest=sha256:cfec2640029faa115bf818fb1bf430d5291df9415fd7f5ead1289e4fa0aaa60c

Observation 1af66703-cc0f-47dd-b29e-8b324fe69872 · inbound

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization cites this paper.

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization Soft Actor-Critic Algorithms and Applications

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:01.334824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:46:27.179341Z digest=sha256:f5172382a5383f3ba95fcbca4988e59d3cc8f5af4226f1e7a6255667542cf18c

Observation 2f4f78f1-c0fa-4fbf-8729-9d23ad3aec25 · inbound

Heterogeneous AAV Logistics Task Allocation: A Reinforcement Learning Enhanced Overlapping Coalition Formation Game Approach cites this paper.

Heterogeneous AAV Logistics Task Allocation: A Reinforcement Learning Enhanced Overlapping Coalition Formation Game Approach Soft Actor-Critic Algorithms and Applications

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:43:45.091597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T17:38:35.614870Z digest=sha256:3a85c30c0cca65e072a92537eb7fe15b3dc45d5d7c49ac3f4bbd007c217c719c

Observation f7fa21a0-6846-4e33-b6cc-81b67887081e · inbound

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient cites this paper.

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient Soft Actor-Critic Algorithms and Applications

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:03:48.614373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T17:34:41.053725Z digest=sha256:20f7f156eb78a81e194eab89786d09166c179ee558726e77005e40dcacff7286

Observation 75d29eb4-16df-438c-b73d-d77b161a34f6 · inbound

Adaptive Reinforcement Learning for Robust Open Quantum System Control: A Multi-Task Framework with Temporal Optimization cites this paper.

Adaptive Reinforcement Learning for Robust Open Quantum System Control: A Multi-Task Framework with Temporal Optimization Soft Actor-Critic Algorithms and Applications

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:03:40.725954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T17:01:28.972549Z digest=sha256:f613055547509ca64347790597ae6bcabc725908c797d9ad80c8490c73f2f3fd

Observation 9642b65c-f81c-4540-b7d6-0277269d4224 · inbound

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning cites this paper.

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:33:31.068975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T06:37:52.331084Z digest=sha256:303bea983a005da71b3eb34a02649ea7c4c4b72ff983cf5d9e80d461f95f0188

Observation 5add10f4-062f-4cc2-a46e-fcc1e5c470c2 · inbound

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning cites this paper.

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T15:37:48.987092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T15:37:48.987092Z digest=sha256:cd11eefb2264c459e41a9d9602f8aa4fb8dd65c4bdc4b548c97fa933bee2c41f

Observation 908e6a21-0cb5-4c9d-93c0-24de49ec733a · inbound

Direct High-Magnetic-Field Coupling to Stripe Order in a Cuprate Superconductor cites this paper.

Direct High-Magnetic-Field Coupling to Stripe Order in a Cuprate Superconductor Soft Actor-Critic Algorithms and Applications

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-06-27T20:41:15.192580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T20:31:38.157391Z digest=sha256:fc0018ef93b41a3c80f785e7d7b9c30fe6e8f9940a61b49853be5472a73d0261

Observation 852670b4-f74e-4d17-bb46-63804936b03f · inbound

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Deep Reinforcement Learning cites this paper.

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Deep Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.106659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T10:56:31.877119Z digest=sha256:0ac7055e2ea535a998a7029c380577f72849bc6f876e8d4f34f9633c704225ad

Observation 709adedb-9564-40d2-9354-47072441e2ad · inbound

Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target cites this paper.

Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target Soft Actor-Critic Algorithms and Applications

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T12:48:11.807231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T08:45:09.342362Z digest=sha256:0eb0504cb8ac9b41530f4c4abd9ff6a24cbebab4af7d033a633d6e31458f416d

Observation 96c91ba2-5833-4094-aa8d-80ab9e5d1b00 · inbound

Time-Slotted Multi-Cluster UAV AirComp with Energy-Awareness: A Pointer Network-Assisted Soft Actor-Critic Learning Framework cites this paper.

Time-Slotted Multi-Cluster UAV AirComp with Energy-Awareness: A Pointer Network-Assisted Soft Actor-Critic Learning Framework Soft Actor-Critic Algorithms and Applications

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:09:01.091259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T22:59:54.788705Z digest=sha256:411439c6a8776a84f39fe1bb2b73cb52ab0e6d340ca5913fa59bcce84b171dec

Observation 199b1bbe-590a-4160-abb5-1ba3ecfa0379 · inbound

Augmenting Game AI with Deep Reinforcement Learning cites this paper.

Augmenting Game AI with Deep Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:59:33.025466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:28:00.929990Z digest=sha256:5786d02f857698d027801de9c35a412884bad1d1518be4a8f1934b5727b16081

Observation 631f69be-bbcb-42b8-8a4f-ffa630c6ebc2 · inbound

ReLaTS: a Reinforcement Learning-based method for dynamically determining the coupling Time Step in multi-scale simulations of self-gravitating systems cites this paper.

ReLaTS: a Reinforcement Learning-based method for dynamically determining the coupling Time Step in multi-scale simulations of self-gravitating systems Soft Actor-Critic Algorithms and Applications

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:49:38.233933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T15:17:55.188215Z digest=sha256:4694b6f314c3c135dc426dfa90b6172beb46ae00c580587b40431eb0666d6ef6

Observation 8fd919cc-c036-4be4-8388-ec4d777416f7 · inbound

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization cites this paper.

Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization Soft Actor-Critic Algorithms and Applications

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:09:37.590741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T14:48:19.683101Z digest=sha256:f67c51fe53fc44b72e4d6f8b489b9fe2f592c434df365196ddd3da49a9ddfb99

Observation f78aa3de-6ca5-43b0-9713-86d047a9f5f5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Soft Actor-Critic Algorithms and Applications

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.698255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:cbcdeb5d5529d6a6da1fdd7004ffafeb2f96c37d007a051d0d0185cca015a549