Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T19:29:21.558661Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 75 inbound Pith citation observations for arXiv:1910.01708.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T19:29:21.558661Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:36:41.833362Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T19:30:07.900798Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f5c71f76-acb0-4258-8d31-5b61f67e85fa · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms An Optimistic Perspective on Offline Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 270e8bfe-f580-4423-b730-4c33b498c8e0 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Exploration by Random Network Distillation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38573401-85c1-45e1-9cf2-e47913c4d24c · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Dopamine: A Research Framework for Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c5238db0-d2c2-4881-8326-65a18ad5bebf · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Distributional Reinforcement Learning with Quantile Regression
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4037aa9b-b0a9-4256-8a13-c2ca701c4aa1 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Off-policy deep reinforcement learning without exploration
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0aba04cd-9a8c-4beb-aaf0-e641fb6a70f4 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Horizon: Facebook's Open Source Applied Reinforcement Learning Platform
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0645ee62-c2cb-4bf8-8b04-7a84ec627f0a · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Stable function approximation in dynamic programming
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a53f4dce-7733-4839-9ccb-0d0d27b6c8e8 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Rainbow: Combining Improvements in Deep Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8ecc38f3-7f8c-4fed-88d4-79f0c81936f2 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d842aee0-fa97-44ef-8a4c-7f3a8b62cdf3 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Adam: A Method for Stochastic Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2898cbef-9b95-4710-abf4-7ae54f40de23 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9d6aff67-f8ff-41f2-8fe6-f18b8d49dc25 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Continuous control with deep reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce67a85b-069e-40d6-844f-a71c3285e649 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Safe Policy Improvement with an Estimated Baseline Policy
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb6f1ee4-bccf-47ae-b827-6e3fb5c8789d · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Dropout: a simple way to prevent neural networks from overfitting
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 50465d77-4b46-4af1-b0c0-b30612ec889d · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Issues in using function approximation for reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c80d18f-b833-4539-9c3f-26ef7f04f701 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Deep reinforcement learning with double q- learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b4b4e93e-f71b-4152-a014-c9282939f8cc · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c231a451-7b55-4a9b-ac67-165dff134b61 · outbound
Benchmarking Batch Deep Reinforcement Learning Algorithms Table 1: Hyper-parameters used by each network
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e568771c-86df-498e-9c1c-d17c6d0e787f · inbound
Analyzing Adversarial Inputs in Deep Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10774b93-9586-4dc5-adb2-f478da525fe6 · inbound
An Investigation of Offline Reinforcement Learning in Factorisable Action Spaces Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db3688e-30cd-44f8-b1ed-664c158916af · inbound
Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402a4b1c-c5b2-4101-9cfa-b4172a00738e · inbound
TEA: Trajectory Encoding Augmentation for Robust and Transferable Policies in Offline Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601e906b-7322-4f7b-b574-398969d00804 · inbound
A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f35f8a-f1fd-4bb7-b509-1ae6ea2e49c5 · inbound
Action Mapping for Reinforcement Learning in Continuous Environments with Constraints Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3816790-3ba9-42c1-b2dd-5b96f7001b27 · inbound
Effective Reward Specification in Deep Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 257
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d574a47f-0de8-4f86-82fc-c3e688622d7b · inbound
From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc362bd-6f14-461d-b3f2-e4e01432b57c · inbound
Deep Reinforcement Learning for Scalable Multiagent Spacecraft Inspection Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · inbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a019047-0e98-42fb-8f3e-472e92eea9e5 · inbound
Offline Safe Reinforcement Learning Using Trajectory Classification Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51d4465-e89a-44f6-9583-746d3397ec41 · inbound
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b48b65a-0ac4-4b9e-a035-989acc42a4c1 · inbound
Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512aa79d-cd0c-4e7e-a916-740147faf020 · inbound
Scalable Safe Multi-Agent Reinforcement Learning for Multi-Agent System Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4b5a4d-393b-48bc-b55e-c00e0b9c856a · inbound
DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60264791-448b-420b-b7ae-e530331b3bdc · inbound
Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceac96ca-22cb-4c6f-af82-1ec1d217ac9c · inbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480eab5c-7607-40a5-a969-021318ced5d9 · inbound
TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d7b077-9df4-4c27-9054-1848b8900735 · inbound
Safe Primal-Dual Optimization with a Single Smooth Constraint Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38e2e38-b1b6-4e2c-b8f4-fb894cd11c8c · inbound
Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RL Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba153c4-cc06-4dc8-bf5c-d9d8f999b133 · inbound
Central Path Proximal Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01e1076-39e8-40df-8526-f40b2cf432d8 · inbound
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e781352a-854b-40d1-b5ca-798bcd7c38df · inbound
Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 667f0cfc-f37b-4e6f-90c8-7d41756cd78f · inbound
Joint User Priority and Power Scheduling for QoS-Aware WMMSE Precoding: A Constrained-Actor Attentive-Critic Approach Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9a5b69-4242-44f9-86ca-753cd1795400 · inbound
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d838393-106d-4239-9e0d-dc73312485de · inbound
PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef68a73-05f4-4064-ab42-54a426e0dff8 · inbound
Red-Team Multi-Agent Reinforcement Learning for Emergency Braking Scenario Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3f99c7-1864-46b0-97c3-b30cd9b1c067 · inbound
Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9579f13e-dca8-4ead-b013-1ec71c48ac9b · inbound
A Review On Safe Reinforcement Learning Using Lyapunov and Barrier Functions Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1598ab19-b2c5-43fb-82f7-c67e9faca4b6 · inbound
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb873a62-d66b-4a2e-9b11-428e9547c14e · inbound
Constraint-Aware Reinforcement Learning via Adaptive Action Scaling Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 43ca3ec6-2bf1-47a1-a028-fb0b8ae43eb3 · inbound
A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3685c4cd-4d4c-40e1-b98d-06dcfc111f1c · inbound
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae766d1d-e08d-4457-a588-2b9920ed99b1 · inbound
RVN-Bench: A Benchmark for Reactive Visual Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e75ba4-2f52-4348-90d7-a848a3af5267 · inbound
Fatigue-Aware Learning to Defer via Constrained Optimisation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7cd178e8-a547-4bb8-b0aa-78ebced09910 · inbound
Improving Feasibility via Fast Autoencoder-Based Projections Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59f7449d-b563-4ec3-86cf-cffb7c9d3c12 · inbound
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b751cafb-209d-4cd3-acb2-ce23f9db898e · inbound
CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15b761da-e125-49c1-b1b8-d4b3cea2f781 · inbound
Safe reinforcement learning with online filtering for fatigue-predictive human-robot task planning and allocation in production Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e74a425-c637-49bf-9585-35394d799919 · inbound
CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb937938-e8a1-4f07-9f61-59a2006b624c · inbound
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e779dbca-7a79-428a-a1de-3916ca5ada33 · inbound
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b301ab83-4438-49ac-8af7-1f2ce61b8075 · inbound
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b465eb59-5652-4688-b0b1-4887c2b5b33d · inbound
Semantic State Abstraction Interfaces for LLM-Augmented Portfolio Decisions: Multi-Axis News Decomposition and RL Diagnostics Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bb6e9414-5957-4b52-9384-1c9d22bfa0a2 · inbound
Learned Lyapunov Shielding for Adaptive Control Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 23d99d08-474a-4b3a-ab94-bfc037e87d44 · inbound
Why Does Agentic Safety Fail to Generalize Across Tasks? Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b4d1d4c-fc43-4979-97a0-04b888d5fd60 · inbound
Shaping Zero-Shot Coordination via State Blocking Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e11f91c-3fe9-4a78-8dcc-c908d31193c3 · inbound
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation af15332e-7b74-47fc-9557-f7b69ab34744 · inbound
Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 97bb7f74-5570-410e-99d8-fecde9a46bfc · inbound
Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38f0fb72-863f-40f9-8620-18fe86dc24e1 · inbound
Optimal design of solar-battery hybrid resources considering multi-market participation under weather and price uncertainty Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7d0137ea-c6eb-4324-a1e3-2a6a45929180 · inbound
Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f66ff6a4-cf26-478c-a5b5-3a132ecf6470 · inbound
Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3c9f68c-d675-4e85-898c-3fe9d2c47a5e · inbound
Generative Auto-Bidding with Unified Modeling and Exploration Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b980eb7d-833c-46b2-bd15-e9963f6a5b42 · inbound
Generative OOD-regularized Model-based Policy Optimization Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fc9d92df-4334-4bcd-a42d-310fcb0242fc · inbound
RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b9589d97-1050-4eea-b859-7a76ef6f6d7f · inbound
Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fd9b573-b02a-401f-a5ba-fb9b13bfb157 · inbound
TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7f5b16a5-9a59-40d9-9bca-036bae8ef484 · inbound
CRAX: Fast Safe Reinforcement Learning Benchmarking Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e698c42b-0500-4172-af39-ebb2c0abd390 · inbound
Imagine to Ensure Safety in Hierarchical Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df0a08a4-fc2a-4276-b45d-a5cb9f8001be · inbound
Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e8cd97e5-069e-43df-b886-01e995a3db53 · inbound
Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7a5f7e-4b67-4e41-b1cb-0a6987a21110 · inbound
ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8fa64724-0c3f-4bfb-9582-b0889b6ca1eb · inbound
ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ee01f2c-7765-48d8-9460-9b2f934d0a60 · inbound
PPO-EAL: Exact Augmented Lagrangian Proximal Policy Optimization for Safe Robotic Control Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 637a075c-bb27-4f0c-89b6-dc5bcad91686 · inbound
Safe Online Learning via Smooth Safety-Structured Policy Composition Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94e178dc-7981-4ea0-8e55-5c4db24dcc9a · inbound
OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44759c31-c891-4ed8-af44-f173415ead9c · inbound
Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c410a26a-eb3d-4913-aeb1-590e5959ffd1 · inbound
Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 761f3ebe-82a2-4654-98b5-51709527c71f · inbound
Fused Constrained Policy Reuse Optimization for Wireless Resource Allocation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 492df8cf-689d-419a-b69c-0b7cfeb71747 · inbound
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2763f45d-b120-4eb2-b30d-03bcd913fe23 · inbound
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a0de04b-cf77-4da2-b511-dfac7aaf53c7 · inbound
EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a5b947-86cf-4c11-b9e4-2c2b6c079aba · inbound
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f95649-afec-4334-8810-b788610e90f5 · inbound
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.