Pith. sign in

Paper Citation Record · LEDGER

Soft Actor-Critic for Discrete Action Settings

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:1910.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.07207 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:52:14.128108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:05.986446Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 67f86b76-e5df-433a-9d5b-4d957e1c4d7f · inbound

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers cites this paper.

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers Soft Actor-Critic for Discrete Action Settings

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:14.128108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:14.128108Z digest=sha256:0aafe014d08f420fcd3ff738f965c89eba4f51ad80f913c8eda7694a9d101b02

Observation ea992a80-c1bd-4a3e-9791-c4d206ad2e72 · inbound

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed cites this paper.

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed Soft Actor-Critic for Discrete Action Settings

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T10:52:50.944478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:52:50.944478Z digest=sha256:c5664145629138d39b07b56b922615491da6653155a49635f397fb12cadf6b31

Observation 6202d432-773c-4208-affb-58f1ba6d27dc · inbound

Dream to Drive with Predictive Individual World Model cites this paper.

Dream to Drive with Predictive Individual World Model Soft Actor-Critic for Discrete Action Settings

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T11:08:53.322150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:08:53.322150Z digest=sha256:f4115d4215b386899eb4ecb1d702db286f250b91be986a1aeb6ade96d4a4a4ca

Observation ebce658c-c843-4fdb-a824-5f93eff0308e · inbound

Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning cites this paper.

Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T21:16:56.312627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:16:56.312627Z digest=sha256:e3a3224abdd4a43aaccdd5ba2bf42053754c813d81dfd7a64b39e77d85fe7ef6

Observation 297a5d88-0833-4fa6-9b98-c16ed13bda4e · inbound

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning cites this paper.

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T21:39:25.525912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:39:25.525912Z digest=sha256:1433ed9ca2f44ff21c0e4d2f7db1b3faf50bb2a0802ed87ab441a29fe9df52fa

Observation 5e975619-5ca9-4778-966f-6b1b82faa63c · inbound

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training cites this paper.

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training Soft Actor-Critic for Discrete Action Settings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:10.179732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:10.179732Z digest=sha256:4c0bf225387301533d71a5b1d5b43ee508df98dead6d9b0a1f6ad1c00b0cd1f6

Observation c81fe141-3c17-4314-a79f-5993e0643127 · inbound

Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints cites this paper.

Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints Soft Actor-Critic for Discrete Action Settings

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:47:14.935311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:14.935311Z digest=sha256:fd86f7e48c48d9e17c6fe60a52b624852770d3d739136c634e3342b6373d60fd

Observation 7b6c6c6b-ad19-4b2e-8a59-d6bf083e2595 · inbound

Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections cites this paper.

Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections Soft Actor-Critic for Discrete Action Settings

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.797853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:14.797853Z digest=sha256:908d8146ce6b2c054da5c0b4b61a3b346faf7b87216568e6d7847f77bb83efcd

Observation 7a9a86bb-cec3-4ae3-a691-68ba02f6f199 · inbound

Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits cites this paper.

Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits Soft Actor-Critic for Discrete Action Settings

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:19.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:19.534851Z digest=sha256:d757ae57210347b20b24d8d1b964830fa168bce608fc156e95e23f699f3f5b02

Observation af61d2fc-ea81-4b3b-a899-4dd8e4a604de · inbound

Learning To Communicate Over An Unknown Shared Network cites this paper.

Learning To Communicate Over An Unknown Shared Network Soft Actor-Critic for Discrete Action Settings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:05.218895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:05.218895Z digest=sha256:19f92c159800d42e9b8116161377ee7882892ec39daf09d8c745088453771299

Observation 9de1f464-3e53-4755-8767-6d82f0270712 · inbound

Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance cites this paper.

Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance Soft Actor-Critic for Discrete Action Settings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:26.517932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:26.517932Z digest=sha256:e1b50c6cede96615f0f2c5a44a073f8dd14f6b7d0f5d4b8efe079a1a5db634b1

Observation 50fb1378-d1f6-4cb1-9e71-21cdd1439884 · inbound

Relative Entropy Pathwise Policy Optimization cites this paper.

Relative Entropy Pathwise Policy Optimization Soft Actor-Critic for Discrete Action Settings

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T04:22:04.036090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T04:19:15.018380Z digest=sha256:3daf379ddd75d0c0c3d9b61d32d10cf959a678f6970a5aa227c6d0efa8c61fec

Observation 8a37a432-c78c-4615-a09a-5be87034d878 · inbound

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing cites this paper.

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing Soft Actor-Critic for Discrete Action Settings

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:29.206067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:29.206067Z digest=sha256:69190e79cb08f54defea3fb089848e7a57524473b5ca0f85faaf4afd5f5ccc3c

Observation 45a43ee0-7ece-4498-97a2-fe9732ceeb15 · inbound

DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning cites this paper.

DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:03.043402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:03.043402Z digest=sha256:883c00f704d119fe269cdd48d99ec2174dcf1e2e2e3706ee001aaf700759f71c

Observation 5ea9f523-b5a7-47d8-9b2a-551b021f2ce8 · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Soft Actor-Critic for Discrete Action Settings

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:06:39.804243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:d7926891eca8220fd2e2a6bcab60bf40b35df1946abeced240e9813c59117c4d

Observation fd073855-f9c5-42f0-b93a-14c5a3161913 · inbound

Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing cites this paper.

Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:59.079116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:59.079116Z digest=sha256:367ad29a73f89494605e6e2a03c78769612de461646e25afcacd79c492d383e7

Observation 6cd3e863-f66f-4f17-a3a8-b8df328c222a · inbound

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments cites this paper.

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments Soft Actor-Critic for Discrete Action Settings

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:10:33.889159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T01:10:12.195443Z digest=sha256:5bc6da1a59318f4205725351e02c4877289ba4e63546a83baade27930daa0873

Observation 16501943-0d04-4c7d-ad33-cb1a8e2a7d5e · inbound

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability cites this paper.

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability Soft Actor-Critic for Discrete Action Settings

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:20:11.613626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:19:24.841090Z digest=sha256:effb4f435fd9a4329d7d31b11aa78a4a89ae14cbfc7045c120f384a71df5ddd4

Observation b59e7892-1632-4b36-9df7-20360aefeeee · inbound

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding cites this paper.

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding Soft Actor-Critic for Discrete Action Settings

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T14:51:03.626575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:51:03.626575Z digest=sha256:eb4a9a28b91bbf8e12cccae5f35c2968cbf471b4faf3613173c77396a2ec11df

Observation 3f2a1f31-9eaf-4298-bd9e-47f1e69b87a8 · inbound

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning cites this paper.

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:16:12.944273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-22T07:15:32.024198Z digest=sha256:1ff73b75dc35957b85cc7b165c6fa031fbc720e410c15cd236a4a5c3587b5861

Observation 4ce9cad1-9684-465b-a774-bee3b9e4f106 · inbound

Retry Policy Gradients in Continuous Action Spaces cites this paper.

Retry Policy Gradients in Continuous Action Spaces Soft Actor-Critic for Discrete Action Settings

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.023856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:cf3390b6ed0913e7d24eaa42635cb8d88eff5cbb06a9d7646b1e98b2c9f54490

Observation caf878c9-6fa1-4d62-9296-9b5c8914068b · inbound

Your GFlowNet Secretly Learns an Optimal Transport Plan cites this paper.

Your GFlowNet Secretly Learns an Optimal Transport Plan Soft Actor-Critic for Discrete Action Settings

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.527589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T02:35:08.323804Z digest=sha256:49d8e17bbbc435ca86286dec8f14e602b89ce2a9cd28231b0513f0ad501f0af5

Observation 21e4be6e-28a0-418b-a2a5-a72dab9f0226 · inbound

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection cites this paper.

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.764532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:42:49.740745Z digest=sha256:fc219c02d4fc83b4e264b7e28dc81e7bd3f78ae8c07b5f071d69710fdde5b170

Observation bfa06eaf-6f94-443d-af2f-9c75b1f544cf · inbound

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication cites this paper.

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication Soft Actor-Critic for Discrete Action Settings

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:36.639368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:14:38.297507Z digest=sha256:1471ad60918c4c5eb5733e0b276cac7f33fb6c270004c13c6756e183be2d8642

Observation 1ac1056f-70e0-4e87-92f9-c9aba62a0250 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Soft Actor-Critic for Discrete Action Settings

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.643542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:57c4945967f5e9e1e5b96082547aa227cbf0805c06f30431cc13004a9237ea60

Observation 18408309-491c-402f-9e55-7c8bf0147e3e · inbound

FactorLibrary: From Polynomials to Circuits via Recursive Subgoals cites this paper.

FactorLibrary: From Polynomials to Circuits via Recursive Subgoals Soft Actor-Critic for Discrete Action Settings

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:05.988216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T21:12:22.718409Z digest=sha256:37e71825f0e77ef518b09d1f0ea299cb717eee1d5bee5cab7042b2f766c0e224

Observation cc54dc00-345d-4228-9582-bdbcec8f327d · inbound

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning cites this paper.

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:05.935795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:05.935795Z digest=sha256:4f79d8b6eb7ba9fafd876a258ed99a40ad2a0a12afed3cb48eb1a48df89f2a93

Observation 406d8430-2d0c-4105-9d0e-a070b86eef4b · inbound

Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening cites this paper.

Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening Soft Actor-Critic for Discrete Action Settings

Reference 345

Resolution
unresolved
no resolver link, observed 2026-08-01T06:15:18.525217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:15:18.525217Z digest=sha256:aa3b22866a4ca3335a237492324a6acde22826b8654dc6772bb6a2d7c8fdec2b

Observation 5d729623-6732-4a19-b31b-0a77d438c5f2 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic for Discrete Action Settings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.255389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.255389Z digest=sha256:13b54b7b0ed19a46cea2f124c302cf03c3a1828d14dae81abfbc5fd0157bba34

Observation 641735f2-9b57-48e6-98be-efa0f519e538 · inbound

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition cites this paper.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.061953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.061953Z digest=sha256:7d877c8fdc8fe2d7d1189105784313e62f1de09d1de479b88ad9303e78a9facd