Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Batch Deep Reinforcement Learning Algorithms

As of 23 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 75 inbound Pith citation observations for arXiv:1910.01708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.01708 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T19:29:21.558661Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:36:41.833362Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:30:07.900798Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact10
  • verified fuzzy6
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5c71f76-acb0-4258-8d31-5b61f67e85fa · outbound

This paper cites An Optimistic Perspective on Offline Reinforcement Learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms An Optimistic Perspective on Offline Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.590209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:e84d98af3e795681b68a36a547c0c0b60ee1cfd87808f40ddb1521ff88e1d822

Observation 270e8bfe-f580-4423-b730-4c33b498c8e0 · outbound

This paper cites Exploration by Random Network Distillation.

Benchmarking Batch Deep Reinforcement Learning Algorithms Exploration by Random Network Distillation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.599744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:4a83c800da1787291dfa1d0c0cbea0559fa5b83bddf6e6aa9e1682ed69e662dd

Observation 38573401-85c1-45e1-9cf2-e47913c4d24c · outbound

This paper cites Dopamine: A Research Framework for Deep Reinforcement Learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms Dopamine: A Research Framework for Deep Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.611689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:df213698dccf5555d2dca59faea1a1c430235f9f351d3e82871dfaa825abf528

Observation c5238db0-d2c2-4881-8326-65a18ad5bebf · outbound

This paper cites Distributional Reinforcement Learning with Quantile Regression.

Benchmarking Batch Deep Reinforcement Learning Algorithms Distributional Reinforcement Learning with Quantile Regression

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.618498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:72e16ba2bc003f248463df721da7b32a6291266256f8a3535e579fe55b378b4f

Observation 4037aa9b-b0a9-4256-8a13-c2ca701c4aa1 · outbound

This paper cites Off-policy deep reinforcement learning without exploration.

Benchmarking Batch Deep Reinforcement Learning Algorithms Off-policy deep reinforcement learning without exploration

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.666629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:c7efd39a4289a216f1a28b42fdd5d83e4f241c4128a34cf8537f2a2186495e68

Observation 0aba04cd-9a8c-4beb-aaf0-e641fb6a70f4 · outbound

This paper cites Horizon: Facebook's Open Source Applied Reinforcement Learning Platform.

Benchmarking Batch Deep Reinforcement Learning Algorithms Horizon: Facebook's Open Source Applied Reinforcement Learning Platform

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.625100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:808252051919d0afd568b1e4ecc00acecd1776e5e354667d4ce6ebe0f7ff94a7

Observation 0645ee62-c2cb-4bf8-8b04-7a84ec627f0a · outbound

This paper cites Stable function approximation in dynamic programming.

Benchmarking Batch Deep Reinforcement Learning Algorithms Stable function approximation in dynamic programming

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.674078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:f6b9a2c64f3ec69b9644ab7ef23937bb8ba78b7467d1e3124fcc30b0381f1782

Observation a53f4dce-7733-4839-9ccb-0d0d27b6c8e8 · outbound

This paper cites Rainbow: Combining Improvements in Deep Reinforcement Learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms Rainbow: Combining Improvements in Deep Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.631751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:2dd3cd413ffc7a9c6a4217425c24b7a1ffbbd1fe2556c1e77d1316a928d68653

Observation 8ecc38f3-7f8c-4fed-88d4-79f0c81936f2 · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Benchmarking Batch Deep Reinforcement Learning Algorithms Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.637970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:23b8af13b7b62b40d0e6453e52ffdf7c7923f812e888705020478bf37e6bfe47

Observation d842aee0-fa97-44ef-8a4c-7f3a8b62cdf3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Benchmarking Batch Deep Reinforcement Learning Algorithms Adam: A Method for Stochastic Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.644578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:ac3d14c9db6a21905931f773736a47c04581a34e6b6b6a68abd0567c3482466f

Observation 2898cbef-9b95-4710-abf4-7ae54f40de23 · outbound

This paper cites Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction.

Benchmarking Batch Deep Reinforcement Learning Algorithms Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T19:29:21.650917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:f5d3c207d545aeaf9afb741c0847c97552efe6480eabfc83f19a6284c994599a

Observation 9d6aff67-f8ff-41f2-8fe6-f18b8d49dc25 · outbound

This paper cites Continuous control with deep reinforcement learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms Continuous control with deep reinforcement learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:29:21.656717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:3f39baca56d58f9ee9efe49438d7b6e83392dc70a91dcf1c2aac277fac1a815d

Observation ce67a85b-069e-40d6-844f-a71c3285e649 · outbound

This paper cites Safe Policy Improvement with an Estimated Baseline Policy.

Benchmarking Batch Deep Reinforcement Learning Algorithms Safe Policy Improvement with an Estimated Baseline Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.662602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:e83133ee4bb4f143e76ddc6758d6fd4f3eb8df78b34b5120c40bacfc0839227a

Observation cb6f1ee4-bccf-47ae-b827-6e3fb5c8789d · outbound

This paper cites Dropout: a simple way to prevent neural networks from overfitting.

Benchmarking Batch Deep Reinforcement Learning Algorithms Dropout: a simple way to prevent neural networks from overfitting

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.681274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:e3ddef9b53bec24724233f9647a26e0c7e5bed6dcbf38f9b1114b8f821fbc9a3

Observation 50465d77-4b46-4af1-b0c0-b30612ec889d · outbound

This paper cites Issues in using function approximation for reinforcement learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms Issues in using function approximation for reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.684791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:9c5212e697e13e1faf216cf809bbbf1d3ad1b3bfc053d7f20f0b6281f25c7b00

Observation 6c80d18f-b833-4539-9c3f-26ef7f04f701 · outbound

This paper cites Deep reinforcement learning with double q- learning.

Benchmarking Batch Deep Reinforcement Learning Algorithms Deep reinforcement learning with double q- learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.688632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:387dd9366f74a2e699d6bc862259a867fe9219b89fa41df3a80608e23d537e24

Observation b4b4e93e-f71b-4152-a014-c9282939f8cc · outbound

This paper cites an unresolved cited work.

Benchmarking Batch Deep Reinforcement Learning Algorithms Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:29:21.670458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:686ff58c47ab30d3cae44302e10e4d04c2ef94b68123b1a206745664a309c84f

Observation c231a451-7b55-4a9b-ac67-165dff134b61 · outbound

This paper cites Table 1: Hyper-parameters used by each network.

Benchmarking Batch Deep Reinforcement Learning Algorithms Table 1: Hyper-parameters used by each network

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:29:21.677555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T19:29:21.558661Z digest=sha256:c4cb7a922d1d17b925bf37cc187a37ec558993593e4076f32e5443443b93d60f

Pith citing papers

Observation e568771c-86df-498e-9c1c-d17c6d0e787f · inbound

Analyzing Adversarial Inputs in Deep Reinforcement Learning cites this paper.

Analyzing Adversarial Inputs in Deep Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:43:50.366105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T03:40:04.265426Z digest=sha256:9827f6b50b4a8b89f80a6de7e89cb0ac79de4f145e77a8ad67925853e56dfafe

Observation 10774b93-9586-4dc5-adb2-f478da525fe6 · inbound

An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces cites this paper.

An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:03:10.423609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:03:10.423609Z digest=sha256:9ab239938a9c8b55b7ca54c79bcf0dce6b7a7e836f8d23055ed010178ba96211

Observation 4db3688e-30cd-44f8-b1ed-664c158916af · inbound

Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium cites this paper.

Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:44:45.503266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:44:45.503266Z digest=sha256:1af14a09a141afc8a83d208bceb91a915559806c872f837b68f77f3061bd642e

Observation 402a4b1c-c5b2-4101-9cfa-b4172a00738e · inbound

TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning cites this paper.

TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:32:43.004752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:32:43.004752Z digest=sha256:d1b934307de7ce329847032b092db42d846ead43f967954ca9c9a93c2d0f483b

Observation 601e906b-7322-4f7b-b574-398969d00804 · inbound

A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks cites this paper.

A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:46:47.761953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:46:47.761953Z digest=sha256:29dd733742f44a72fa0779896df55358b7cd2decd0c6c6bdcaaccf6918207112

Observation 96f35f8a-f1fd-4bb7-b509-1ae6ea2e49c5 · inbound

Action Mapping for Reinforcement Learning in Continuous Environments with Constraints cites this paper.

Action Mapping for Reinforcement Learning in Continuous Environments with Constraints Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:39:11.960949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:39:11.960949Z digest=sha256:41603e10c57e30e9e59a3c01bf89429214970602b9c4d2bc937510b9534f7b7b

Observation c3816790-3ba9-42c1-b2dd-5b96f7001b27 · inbound

Effective Reward Specification in Deep Reinforcement Learning cites this paper.

Effective Reward Specification in Deep Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:55.310027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:55.310027Z digest=sha256:e439bf29f11f7ddc84cb538c1e6b9074914d56795ce533627702f9f810a18e4a

Observation d574a47f-0de8-4f86-82fc-c3e688622d7b · inbound

From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning cites this paper.

From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:14.308610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:14.308610Z digest=sha256:128fda7219f6d047a72a7f10c11246ff8a5e81451456373f9725b100b3560396

Observation 5fc362bd-6f14-461d-b3f2-e4e01432b57c · inbound

Deep Reinforcement Learning for Scalable Multiagent Spacecraft Inspection cites this paper.

Deep Reinforcement Learning for Scalable Multiagent Spacecraft Inspection Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:54.063493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:54.063493Z digest=sha256:0f568010bc4c45c2c1be0abf15b4f8c669beb3991a031e66caa5abf954335a48

Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · inbound

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation cites this paper.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.465505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.465505Z digest=sha256:f00fbe3618b43dcf072c15d89fe8139da1e260f62b59c4977e9c4302638cf137

Observation 4a019047-0e98-42fb-8f3e-472e92eea9e5 · inbound

Offline Safe Reinforcement Learning Using Trajectory Classification cites this paper.

Offline Safe Reinforcement Learning Using Trajectory Classification Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:06.696490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:06.696490Z digest=sha256:b2b16cbf6924e01064cd2089462808a5b3b6ccaf611f8718ab0ee2e3de939d71

Observation c51d4465-e89a-44f6-9583-746d3397ec41 · inbound

Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning cites this paper.

Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:23:19.687726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:23:19.687726Z digest=sha256:da21af50514298550bae3fa1cf215740e2549b3725d09c67bc85b6853e0d182f

Observation 9b48b65a-0ac4-4b9e-a035-989acc42a4c1 · inbound

Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind cites this paper.

Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:04.057723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:04.057723Z digest=sha256:aa33cb92390e8aa02a2b16e98d8c68dd5619bf8e0998dc8970ab8d4cf0801313

Observation 512aa79d-cd0c-4e7e-a916-740147faf020 · inbound

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System cites this paper.

Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T15:42:33.274421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:42:33.274421Z digest=sha256:cfa8acfa85de020c6ddbc36249ceb1c341175a8c7433f79b86ed4730d52714e3

Observation db4b5a4d-393b-48bc-b55e-c00e0b9c856a · inbound

DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems cites this paper.

DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T00:48:32.560569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:48:32.560569Z digest=sha256:c14a1a622a1b739a8c61127f6fd955d58ceaf2ad37a84305ba4a452d7aba807f

Observation 60264791-448b-420b-b7ae-e530331b3bdc · inbound

Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation cites this paper.

Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T18:28:22.464527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:28:22.464527Z digest=sha256:e08ca35ffba90c33189b2a2c27f0d0b79f66b3330c5aa4cbb61a292958f4d18a

Observation ceac96ca-22cb-4c6f-af82-1ec1d217ac9c · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.091162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.091162Z digest=sha256:a0d99ee749b9864973fe279b9f2e594a4ff3b5b1df88878e019010cd5bcccbac

Observation 480eab5c-7607-40a5-a969-021318ced5d9 · inbound

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels cites this paper.

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:36:41.833362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:36:41.833362Z digest=sha256:a045fb463494eda9fd7d1aba7d9bbdba415469421a93d7619c05c0e80dcab89d

Observation 97d7b077-9df4-4c27-9054-1848b8900735 · inbound

Safe Primal-Dual Optimization with a Single Smooth Constraint cites this paper.

Safe Primal-Dual Optimization with a Single Smooth Constraint Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:48:26.219026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:48:26.219026Z digest=sha256:e3f970c4b35156f479ff80e299b2a1b95b37b10e3e73290a541ec2fc1a8bb457

Observation f38e2e38-b1b6-4e2c-b8f4-fb894cd11c8c · inbound

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL cites this paper.

Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:02.667693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:02.667693Z digest=sha256:d77660f14501abea3443119c662b0be61ded4910e1a4382eba60cd4c4949ae1b

Observation 6ba153c4-cc06-4dc8-bf5c-d9d8f999b133 · inbound

Central Path Proximal Policy Optimization cites this paper.

Central Path Proximal Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:08.990567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:07:08.990567Z digest=sha256:8dddad479b3d0ab931e94088e99f542d642ca29fc454f01dd8c839857395f7d6

Observation c01e1076-39e8-40df-8526-f40b2cf432d8 · inbound

Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL cites this paper.

Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:15:39.426563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:15:39.426563Z digest=sha256:8632693d2dbf9faf5fd63cd0753ffe8b9f19b130af0235ce58318f07156d2f6a

Observation e781352a-854b-40d1-b5ca-798bcd7c38df · inbound

Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism cites this paper.

Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:05.554987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:02:05.554987Z digest=sha256:658e9cbf2e99f471a6bd663388bd39cbff30dc7bea7529b65576a5d59d2ddfcf

Observation 667f0cfc-f37b-4e6f-90c8-7d41756cd78f · inbound

Joint User Priority and Power Scheduling for QoS-Aware WMMSE Precoding: A Constrained-Actor Attentive-Critic Approach cites this paper.

Joint User Priority and Power Scheduling for QoS-Aware WMMSE Precoding: A Constrained-Actor Attentive-Critic Approach Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:48.290281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:48.290281Z digest=sha256:1312b5d3f7c8370ea123df89fc3f98e282a63517d7fd139b72868b9f5ee09476

Observation 5e9a5b69-4242-44f9-86ca-753cd1795400 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.086772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.086772Z digest=sha256:66e97019450022c5e535d4bd4ea0080864729d71130158fe4dc0fbb6e871105e

Observation 7d838393-106d-4239-9e0d-dc73312485de · inbound

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning cites this paper.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.412845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.412845Z digest=sha256:cb238245bf2a2ee503d2870dbebc65c8ddb9790224222e01020589557c1d7b30

Observation 2ef68a73-05f4-4064-ab42-54a426e0dff8 · inbound

Red-Team Multi-Agent Reinforcement Learning for Emergency Braking Scenario cites this paper.

Red-Team Multi-Agent Reinforcement Learning for Emergency Braking Scenario Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:56.550308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:56.550308Z digest=sha256:b3dc2d13d36d79a6530a947cdeda64e95f359c33349ed501f535758078e03438

Observation 9e3f99c7-1864-46b0-97c3-b30cd9b1c067 · inbound

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction cites this paper.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:21.442854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:21.442854Z digest=sha256:335eb1c553d4101393ca384316d10e0f3aa22217a752df42de31a2325ab85666

Observation 9579f13e-dca8-4ead-b013-1ec71c48ac9b · inbound

A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions cites this paper.

A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:41:53.789641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T22:37:32.388931Z digest=sha256:5d5a68f8a30d53c6bbd2c7d060c17b7b5204056b73aef52d01e98fceba5541bf

Observation 1598ab19-b2c5-43fb-82f7-c67e9faca4b6 · inbound

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning cites this paper.

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:35:33.289482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:35:33.289482Z digest=sha256:14031fa1b667ab486abc20d84b658b16cac071d45b09f3c5673b90d9a519caa9

Observation fb873a62-d66b-4a2e-9b11-428e9547c14e · inbound

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling cites this paper.

Constraint-Aware Reinforcement Learning via Adaptive Action Scaling Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:41:03.345959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T07:36:28.841586Z digest=sha256:085ef155db931f93b0d1d5ea4fa3f5ef6181433b593b2fa73c2f38d3db909d66

Observation 43ca3ec6-2bf1-47a1-a028-fb0b8ae43eb3 · inbound

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach cites this paper.

A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:52:55.274637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:52:55.274637Z digest=sha256:0e6c779cdb4f849daae6dbd0046d7876601bb0c3dfbe8d3a6426c58d5310a8d6

Observation 3685c4cd-4d4c-40e1-b98d-06dcfc111f1c · inbound

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents cites this paper.

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:51:25.963446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T10:49:35.846590Z digest=sha256:0175a5ec100b471dc10366565fc27299b0dea5280462a315d5d6f7bb1563f7e4

Observation ae766d1d-e08d-4457-a588-2b9920ed99b1 · inbound

RVN-Bench: A Benchmark for Reactive Visual Navigation cites this paper.

RVN-Bench: A Benchmark for Reactive Visual Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:03:43.704845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:03:43.704845Z digest=sha256:47f3024a8080ee6cbaa8a0afcb11352e53cf4fd7d09b3f8fbec427b524a4cbb8

Observation b2e75ba4-2f52-4348-90d7-a848a3af5267 · inbound

Fatigue-Aware Learning to Defer via Constrained Optimisation cites this paper.

Fatigue-Aware Learning to Defer via Constrained Optimisation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T22:40:32.912158Z digest=sha256:a5532d0276fa16bb7cedb87795f8e3932426e773d225fb3badf9ae9a92ba1b81

Observation 7cd178e8-a547-4bb8-b0aa-78ebced09910 · inbound

Improving Feasibility via Fast Autoencoder-Based Projections cites this paper.

Improving Feasibility via Fast Autoencoder-Based Projections Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T19:29:59.328674Z digest=sha256:7f6b84ed2443471c9f27960b9c1e45bbbb21e212308d229c0ef8f998580576a0

Observation 59f7449d-b563-4ec3-86cf-cffb7c9d3c12 · inbound

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems cites this paper.

Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-13T10:37:39.719847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:37:39.719847Z digest=sha256:4df7a3d3047b48c609a620625581ba9e1156ea31f20c83067c9343e9de2b7059

Observation b751cafb-209d-4cd3-acb2-ce23f9db898e · inbound

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection cites this paper.

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:34:10.942429Z digest=sha256:c74af2f3c912fb48ec5acd287370c0799fead5fa34dd41c702508401eddc107e

Observation 15b761da-e125-49c1-b1b8-d4b3cea2f781 · inbound

Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production cites this paper.

Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T14:49:27.919651Z digest=sha256:a61062c9f1338bf912d0c67c7ad16560b102bbeb7afa7896b8ce7f5a8601baff

Observation 8e74a425-c637-49bf-9585-35394d799919 · inbound

CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning cites this paper.

CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T06:30:47.030972Z digest=sha256:608d99750662637b89f54448f38ae81e61f093882cd1b501954dd74a30131946

Observation fb937938-e8a1-4f07-9f61-59a2006b624c · inbound

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning cites this paper.

Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T19:40:42.789401Z digest=sha256:9d773c010b549ff407f424fedc4094243875dd592db88c4f1223cc8fed0a75bc

Observation e779dbca-7a79-428a-a1de-3916ca5ada33 · inbound

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning cites this paper.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T16:34:18.603134Z digest=sha256:94f8493ad781f93e47114f1e0a1e17bfb218fcbda76165c9e8505595d925b890

Observation b301ab83-4438-49ac-8af7-1f2ce61b8075 · inbound

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning cites this paper.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.011373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:ff6c00f247ba0fbb82f5e80127e7b13a8bb4c8d1bf4aa7c87f0a2c28a7ce256e

Observation b465eb59-5652-4688-b0b1-4887c2b5b33d · inbound

Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics cites this paper.

Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:58:24.066300Z digest=sha256:6f61f5b2cc878705d7dcff784b47b127e68228f9f84f452cd7955c7a6bb3d57c

Observation bb6e9414-5957-4b52-9384-1c9d22bfa0a2 · inbound

Learned Lyapunov Shielding for Adaptive Control cites this paper.

Learned Lyapunov Shielding for Adaptive Control Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:02:29.064263Z digest=sha256:7edb4ce88c8f367dd6e0244c4260b835407eb4dd4a48bac2fb1bad626b2e8582

Observation 23d99d08-474a-4b3a-ab94-bfc037e87d44 · inbound

Why Does Agentic Safety Fail to Generalize Across Tasks? cites this paper.

Why Does Agentic Safety Fail to Generalize Across Tasks? Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:55:38.554161Z digest=sha256:c1f527af26d00dd639eb437801cee2d212d99eddb11ef288549756ea3528882b

Observation 6b4d1d4c-fc43-4979-97a0-04b888d5fd60 · inbound

Shaping Zero-Shot Coordination via State Blocking cites this paper.

Shaping Zero-Shot Coordination via State Blocking Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:42:23.260464Z digest=sha256:c60393e940e8b743a4fa4cb7146e4b6e8bb2c071374d6483ac7f0f7784a29f45

Observation 5e11f91c-3fe9-4a78-8dcc-c908d31193c3 · inbound

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs cites this paper.

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:27:12.302109Z digest=sha256:f8b7367b030e0e4be98c842081a6ff7539c58573bd09069e112d3305e936a968

Observation af15332e-7b74-47fc-9557-f7b69ab34744 · inbound

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning cites this paper.

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T07:28:24.455817Z digest=sha256:b8d2b8d60b1696f38adb948d535efca5b408d5c36cd10358a5ab3f3ce69c7aa7

Observation 97bb7f74-5570-410e-99d8-fecde9a46bfc · inbound

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning cites this paper.

Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:49:10.063584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T22:48:55.661356Z digest=sha256:f13199097ef169ba636c7ef747dcb481ca1463c85c5acbdf2ea078818f96f531

Observation 38f0fb72-863f-40f9-8620-18fe86dc24e1 · inbound

Optimal design of solar-battery hybrid resources considering multi-market participation under weather and price uncertainty cites this paper.

Optimal design of solar-battery hybrid resources considering multi-market participation under weather and price uncertainty Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T05:29:01.147758Z digest=sha256:8ab04eaf5a2454437690a9497f7f5d7215ff8f09d624f4c3b43e2ded5d7800b5

Observation 7d0137ea-c6eb-4324-a1e3-2a6a45929180 · inbound

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation cites this paper.

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T04:52:48.491036Z digest=sha256:9e890c78aac7b2ce646704ffd6bd2b2836fc7c9b9e4c91eb5e7936d5673d2a17

Observation f66ff6a4-cf26-478c-a5b5-3a132ecf6470 · inbound

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability cites this paper.

Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:29:21.689617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:50:40.285132Z digest=sha256:89c6feaab9db2882888c859a3f50337ed0582563ebbf26a1a4ab00e7348e1817

Observation c3c9f68c-d675-4e85-898c-3fe9d2c47a5e · inbound

Generative Auto-Bidding with Unified Modeling and Exploration cites this paper.

Generative Auto-Bidding with Unified Modeling and Exploration Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.223920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:36:55.976573Z digest=sha256:80afdf0a569c45d3fe16f584abbdb95abda5e52abdfeac32d74de94077e6bfd0

Observation b980eb7d-833c-46b2-bd15-e9963f6a5b42 · inbound

Generative OOD-regularized Model-based Policy Optimization cites this paper.

Generative OOD-regularized Model-based Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.094477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:42:38.148642Z digest=sha256:5682e05dffac88d56dbad1705f7a3d0f1a10446ac68bb0847eac10e27224ced6

Observation fc9d92df-4334-4bcd-a42d-310fcb0242fc · inbound

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation cites this paper.

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:14:00.560615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T21:10:42.027411Z digest=sha256:870a1977ff07c6d2725b0c14a7f87e961d5661ae5a2d5f062604a6c2f1fe6ddb

Observation b9589d97-1050-4eea-b859-7a76ef6f6d7f · inbound

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents cites this paper.

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:41.598471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T07:34:18.370135Z digest=sha256:a5532a49fe43b161feafb68e7d0e6d1e06a7a7905ea287e204a45cb946d092e6

Observation 9fd9b573-b02a-401f-a5ba-fb9b13bfb157 · inbound

TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning cites this paper.

TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:18:54.937345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T01:54:19.216553Z digest=sha256:a422e47244d2369bfeadf4dedc48ec0beeb3082e658776b4e0134cfedc8e9f41

Observation 7f5b16a5-9a59-40d9-9bca-036bae8ef484 · inbound

CRAX: Fast Safe Reinforcement Learning Benchmarking cites this paper.

CRAX: Fast Safe Reinforcement Learning Benchmarking Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.098710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T18:22:36.891186Z digest=sha256:a53ad0459b8c464fe2ce60568b59cff742cecc86d8c522f2ef304d0b95a4cb0b

Observation e698c42b-0500-4172-af39-ebb2c0abd390 · inbound

Imagine to Ensure Safety in Hierarchical Reinforcement Learning cites this paper.

Imagine to Ensure Safety in Hierarchical Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:59:42.179239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T10:48:11.992980Z digest=sha256:da7a6b972fe8034e42452b9a7c3552388efd9a1c89644f952098764681763b65

Observation df0a08a4-fc2a-4276-b45d-a5cb9f8001be · inbound

Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning cites this paper.

Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:30:07.902034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T21:15:01.199124Z digest=sha256:9f97245b66d8e56fa39419a028a554629329a2a1dda6ec67687aaa312a750e33

Observation e8cd97e5-069e-43df-b886-01e995a3db53 · inbound

Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning cites this paper.

Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T17:22:01.242752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:22:01.242752Z digest=sha256:ac68862df89d4c7c77c55c69be560bfa2dc120dc759eacc38627048366c108ac

Observation 5a7a5f7e-4b67-4e41-b1cb-0a6987a21110 · inbound

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models cites this paper.

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:29:51.836696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T05:10:09.186699Z digest=sha256:dafc3617c6062f64226279ffb00839ed9cc1abad953c98d2e916e52261e9413f

Observation 8fa64724-0c3f-4bfb-9582-b0889b6ca1eb · inbound

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models cites this paper.

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:54:34.851041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T09:51:54.464950Z digest=sha256:9dbd004ba0d207da1d4a91676f83e23cdf9ca482d62c7d67593247133f5c829d

Observation 7ee01f2c-7765-48d8-9460-9b2f934d0a60 · inbound

PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control cites this paper.

PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:13:53.381758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T04:49:22.328169Z digest=sha256:eb5c0951d624e576d6fec66a4948044f12e54c1ce4cf284e47172499f1a0ed6d

Observation 637a075c-bb27-4f0c-89b6-dc5bcad91686 · inbound

Safe Online Learning via Smooth Safety-Structured Policy Composition cites this paper.

Safe Online Learning via Smooth Safety-Structured Policy Composition Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 12

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T09:45:40.506809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:f3944ad6857a50fdf9d9d8c5453d223fe9fd9d8ea79086d7b17c7c7882514b7f

Observation 94e178dc-7981-4ea0-8e55-5c4db24dcc9a · inbound

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation cites this paper.

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:43.040440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:05:51.441647Z digest=sha256:8f1b6184a2ce2a42a11cf19eaceafb5c260812757716acdbe143fa2818779b93

Observation 44759c31-c891-4ed8-af44-f173415ead9c · inbound

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach cites this paper.

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:47:17.440985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T17:43:39.676533Z digest=sha256:e2de8cdaa21086be538059c30acd5f78db2d380a85853288843cea70d2e164c0

Observation c410a26a-eb3d-4913-aeb1-590e5959ffd1 · inbound

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach cites this paper.

Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-03T22:29:00.224900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T22:21:09.950893Z digest=sha256:e8b1bb9b4d0e722e45ba4c75148b17718eb38088a9841211a088b3614cc053bf

Observation 761f3ebe-82a2-4654-98b5-51709527c71f · inbound

Fused Constrained Policy Reuse Optimization for Wireless Resource Allocation cites this paper.

Fused Constrained Policy Reuse Optimization for Wireless Resource Allocation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T02:36:27.505388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:36:27.505388Z digest=sha256:2d86cb54c6543ff854b1697cb2882efeab23b8d16d6c16843e39aa8d8a6e84a6

Observation 492df8cf-689d-419a-b69c-0b7cfeb71747 · inbound

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels cites this paper.

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T00:48:45.749486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:48:45.749486Z digest=sha256:cb7bd7a1cf534749d4aeff3a164f349799ccd33f9564ea659b12fafcd08179e7

Observation 2763f45d-b120-4eb2-b30d-03bcd913fe23 · inbound

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation cites this paper.

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 157

Resolution
unresolved
no resolver link, observed 2026-07-31T14:03:37.225305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:03:37.225305Z digest=sha256:a24c04cd00508a4e25f4d02fcde7e0bbc84a9abf89baa4b1ecc99d61ac4d0c20

Observation 1a0de04b-cf77-4da2-b511-dfac7aaf53c7 · inbound

EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming cites this paper.

EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T21:36:16.177038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:36:16.177038Z digest=sha256:b16277bbcf9d2cde846631af31f86183054d83246dbc29c14878a686196c8824

Observation a7a5b947-86cf-4c11-b9e4-2c2b6c079aba · inbound

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds cites this paper.

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:21:32.228188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:21:32.228188Z digest=sha256:6d7edc3b5148991d5ccc255d0c3a531497d20c87ea37a71bef5450368f164c41

Observation 25f95649-afec-4334-8810-b788610e90f5 · inbound

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning cites this paper.

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:18.373663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:18.373663Z digest=sha256:3ff46d4942661af60c638fb930bc3bdb9d7c0011c9e70ca940c64d3b98fb8b91