Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:56.545141Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.11138.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:56.545141Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0800d1b-1221-4dde-9f03-2fb971020bf8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a360c338-1592-4384-8cbe-68d871306fc0 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained policy optimization, 2017
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a71c32c2-d606-4952-b20a-8bda488a1571 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Markov decision processes, volume 7
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9d7241f-bd17-43b4-910c-2346b9547840 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Policy Optimization via Bayesian World Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b70bcd-f0d2-40d2-b46a-593fd00d3b87 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Robots that interact with humans: a review of safety technologies and standards
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc377145-760d-4253-b4b2-f0818b5ed36c · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Risk-constrained reinforcement learning with percentile risk criteria
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb4e2bc-1b61-424a-980b-d7003ef5fec4 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-Augmented Actor-Critic: Backpropagating through Paths
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 940c8509-2d63-4ecf-9739-8d5275988a23 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Augmented proximal policy optimization for safe reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fa0132e8-2c18-4e22-a256-bf12ef92e5d8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safe RLHF : Safe reinforcement learning from human feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69df94db-b3f8-4867-9d6b-23b7f539f038 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A differentiable physics engine for deep learning in robotics
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e2fba4c-a9f7-4be6-bcaa-350a35f3c12b · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A., Farouk, H., and Mofreh, E
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ec16299-e684-4742-9a86-ad0a989acecd · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 998678a0-0f6a-4dc2-8e83-c9718926bce9 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6da0f8c1-27d8-4aa7-8fb6-e54e96503af9 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Fern \'a ndez, F
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c72b972-cdd9-4fbb-9475-3d4fc74d445a · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Bullet-safety-gym: A framework for constrained reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2bc3246a-e9f7-4e8a-9fbe-5c0dc54e1b5d · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Bhatnagar, S
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation af741505-959b-40aa-86ed-88620819c127 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Personalized robotic control via constrained multi-objective reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd8f973c-9032-4d22-a699-aebf289f3952 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Dojo: A Differentiable Physics Engine for Robotics
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542a055a-db30-4c96-9c4f-51a54f4500b6 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Deep differentiable reinforcement learning and optimal trading
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9df688e-ce39-4a8d-8168-6830c17bb556 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation AI Alignment: A Comprehensive Survey
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb0b150-bbd1-441f-988a-92c973798b93 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1796f7d-c120-452b-b1a1-62d0045df476 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe97cd66-6674-40cc-8114-c2d85b8c6e79 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Aligner: Efficient Alignment by Learning to Correct
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc5257a-0570-4ec5-a019-8bb2b1a35488 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c3ca86f-84c8-49a4-a559-a0d2e92c2df0 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1cbd7cbe-0eea-4de5-b983-a42d7b8e4fea · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Langford, J
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ef741c55-a3d1-444e-949e-cf7bdb033418 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation C., Jain, R., and Nuzzo, P
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 305c086e-6476-4c01-a751-90a7ed9ea8bc · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Reparameterization gradient for non-differentiable models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 376d8cbb-5fde-4d28-bcf4-32f3a7da07f5 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained variational policy optimization for safe reinforcement learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c807d510-2679-498a-a0dd-ba2240b1b7dd · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd0547a-70e9-47d5-906c-7bfe07d721c8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Gradients are Not All You Need
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc24076-c5a0-407f-b5b7-b541570fd9c8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Monte carlo gradient estimation in machine learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 49bd8828-605e-4244-a788-311f55409221 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7acdd6df-4a9a-4ef2-bff5-8f5a322e9b8b · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A focused backpropagation algorithm for temporal pattern recognition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a63348ae-db32-4cce-ab11-1df098d90384 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cbbc98fb-67aa-411f-bb0a-9dd7e69882c8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bf742935-2efa-4048-97d0-0bba24bf954d · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1761bfb-2b04-466e-9cb2-61bcc64c4769 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation M., Smaby, N., and Cutkosky, M
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 712e9719-cd9b-4042-80e5-8bb490350a44 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Training language models to follow instructions with human feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139bfd47-8804-48ff-9cb5-13faf20c5b9b · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3114b496-2514-448b-852b-9f77beb9670a · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 633aab94-066a-433e-8a49-dcfaa483f501 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Barr, A
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 682a5e1f-371a-4840-b23b-9977cfb85249 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation E., Perescu-Popescu, L., and Mastorakis, N
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation efaa3de5-6dd0-45e0-9e11-2630796a9d4a · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a0be68-f149-4d2b-a9ca-6fb830cd6854 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af2260d-6e97-4eb8-8543-0fd2555934bb · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a617aa4-eb65-40b0-9b18-09692828d39e · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trust region policy optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30eb5d28-521f-42c7-9a2d-90b69df97748 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2164f7e4-e74d-4d9f-9cc3-e67046c62bb1 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2a9931-94a1-4178-8e73-2206990cc789 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Mastering the game of go without human knowledge
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4475aa1-6272-45d6-8567-4328e096c74f · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b87282a-bebf-4bd3-99f1-0c0868e88f05 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Responsive safety in reinforcement learning by pid lagrangian methods
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70b62e8d-cedb-4af6-9ad0-e4ab9cc69157 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Simchowitz, M., Zhang, K., and Tedrake, R
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b60e1a78-6ea9-4152-85f0-a6a952922a4b · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a7f6ff-f92e-4263-875e-631c4c3bad3c · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation W., Wang, T., Shang, Y., and Wu, Z
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a320c451-bbe7-4f3a-b23b-6681bab8ddc3 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Development of a humanoid robot control system based on ar-bci and slam navigation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd42231a-daa8-4f4a-9662-c64b4836664a · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39328285-6824-4b56-97bc-6b1ea74a029b · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f83f98c-d743-4be3-8402-224876a928a8 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df34a0f-ee63-4136-a3d1-302b68d056d5 · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 73e89b4b-0b3d-413f-a788-dfa37e890dae · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained update projection approach to safe policy optimization
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f7d4c6dd-6a6a-4971-9ab3-bf5fcb76046f · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Projection-Based Constrained Policy Optimization
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b6b7e47-0add-45be-89ac-407b5fee5bee · outbound
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation First order constrained optimization in policy space
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.