Pith. sign in

Paper Citation Record · LEDGER

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

As of 13 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.11138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11138 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:56.545141Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact4
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0800d1b-1221-4dde-9f03-2fb971020bf8 · outbound

This paper cites write newline.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.249835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.249835Z digest=sha256:adf686dbc8fff88a3779d81f01b8770accf37773a0f0adce38290b698477a09c

Observation a360c338-1592-4384-8cbe-68d871306fc0 · outbound

This paper cites Constrained policy optimization, 2017.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained policy optimization, 2017

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.460464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.255303Z digest=sha256:884922112ac6dbed3c55b320ae51915e4561cc60208ea96ce55925a29415a40d

Observation a71c32c2-d606-4952-b20a-8bda488a1571 · outbound

This paper cites Constrained Markov decision processes, volume 7.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.446656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.260145Z digest=sha256:cd9a600dbb4930be0642bfe90e3e8bf453a7ab452332de97da402f16c1d4a190

Observation d9d7241f-bd17-43b4-910c-2346b9547840 · outbound

This paper cites Constrained Policy Optimization via Bayesian World Models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Policy Optimization via Bayesian World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.263978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.263978Z digest=sha256:4f2a4aecd3cd5c1107dce124505a17dd847a8616db2490f931304a0811cb269b

Observation b5b70bcd-f0d2-40d2-b46a-593fd00d3b87 · outbound

This paper cites Robots that interact with humans: a review of safety technologies and standards.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Robots that interact with humans: a review of safety technologies and standards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.433993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.269227Z digest=sha256:5f85e1eaf6947a573b312f8754f2ed720169b1240524f87549e29a545a380650

Observation cc377145-760d-4253-b4b2-f0818b5ed36c · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Risk-constrained reinforcement learning with percentile risk criteria

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.273249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.273249Z digest=sha256:a2474519f613d8e01c5d85785f667740457eac54b179a8b80c7a73fec4ee8012

Observation feb4e2bc-1b61-424a-980b-d7003ef5fec4 · outbound

This paper cites Model-Augmented Actor-Critic: Backpropagating through Paths.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-Augmented Actor-Critic: Backpropagating through Paths

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.278127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.278127Z digest=sha256:c52e05e7d1b28d3b4527f96f8b1cee2ef5474600dcc890b6ceada86176ad1581

Observation 940c8509-2d63-4ecf-9739-8d5275988a23 · outbound

This paper cites Augmented proximal policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Augmented proximal policy optimization for safe reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.408861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.283465Z digest=sha256:7eeec2bcd7aae5f4e4e2416cbe9a332c90ee94fc1086e8cfc61bc0980932f11b

Observation fa0132e8-2c18-4e22-a256-bf12ef92e5d8 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safe RLHF : Safe reinforcement learning from human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.290448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.290448Z digest=sha256:043a8f9a37646a529dfd35a1c1bcb38eb0f2302e6f23ef9f031d3ee1c25d11c2

Observation 69df94db-b3f8-4867-9d6b-23b7f539f038 · outbound

This paper cites A differentiable physics engine for deep learning in robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A differentiable physics engine for deep learning in robotics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.388846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.296012Z digest=sha256:2eaa5a91ba6c01cd2991d6bd5f7eb7b71119a7d64c407ffe928cc3a58fa7797b

Observation 7e2fba4c-a9f7-4be6-bcaa-350a35f3c12b · outbound

This paper cites A., Farouk, H., and Mofreh, E.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A., Farouk, H., and Mofreh, E

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.376244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.302087Z digest=sha256:04c0a796075c2b8fdc22313e7ce362ced8775afa53056950aa35db2c58dc3679

Observation 7ec16299-e684-4742-9a86-ad0a989acecd · outbound

This paper cites D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.364745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.306751Z digest=sha256:a213d6fe382102d1c74fb9e6b50274263c5c89de91b825ba879137e04ef7d0cd

Observation 998678a0-0f6a-4dc2-8e83-c9718926bce9 · outbound

This paper cites A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System

Reference 13

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.597981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.311323Z digest=sha256:9fc9de5fb6815dd471af2327cb454ff7e568a4ab1563eba0d19c844ea5d882e5

Observation 6da0f8c1-27d8-4aa7-8fb6-e54e96503af9 · outbound

This paper cites and Fern \'a ndez, F.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Fern \'a ndez, F

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.318141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.318141Z digest=sha256:fd662d2fff3b1d5ac75190cb582738e16d357b86131b22dfc992f9138c746f6d

Observation 2c72b972-cdd9-4fbb-9475-3d4fc74d445a · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.345538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.322058Z digest=sha256:2bde1d372d206bbae1c2810ce68af45b829da160849c359ab0c35f20316a5f81

Observation 2bc3246a-e9f7-4e8a-9fbe-5c0dc54e1b5d · outbound

This paper cites and Bhatnagar, S.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Bhatnagar, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.333323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.327290Z digest=sha256:bc6c03ad65dbb9ab0ba0faf5d25b315d936aaedca59678ca05d9d4b77c36b554

Observation af741505-959b-40aa-86ed-88620819c127 · outbound

This paper cites Personalized robotic control via constrained multi-objective reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Personalized robotic control via constrained multi-objective reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.314490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.331700Z digest=sha256:01e17e3949be621d703a6332430ec927ef313ea6c293deae68600839615a8bbf

Observation dd8f973c-9032-4d22-a699-aebf289f3952 · outbound

This paper cites Dojo: A Differentiable Physics Engine for Robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Dojo: A Differentiable Physics Engine for Robotics

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.336180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.336180Z digest=sha256:16956207edc83dacd036aabb9a4ccc0754ef0428b99777d0f7e542a7860e134c

Observation 542a055a-db30-4c96-9c4f-51a54f4500b6 · outbound

This paper cites Deep differentiable reinforcement learning and optimal trading.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Deep differentiable reinforcement learning and optimal trading

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.287611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.341108Z digest=sha256:0091bf99c7a207bc9c0ce016704e69249c5501934db259bec7387f7a0bab1de6

Observation d9df688e-ce39-4a8d-8168-6830c17bb556 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation AI Alignment: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.344908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.344908Z digest=sha256:0e3c06488da9163f6e8f694ac45cfdebd3dcb64e06aa3ba44f0a023059ea9fdc

Observation bbb0b150-bbd1-441f-988a-92c973798b93 · outbound

This paper cites Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.349062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.349062Z digest=sha256:ec780d8d83efb90c3fd84df48dce9ba23e9ab49cc4a38baa117c8bf546965fff

Observation b1796f7d-c120-452b-b1a1-62d0045df476 · outbound

This paper cites OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.353163Z digest=sha256:529f2a08750f97242f49ec7c195bdb92e054a215fb0e9cd650de1031e1f29c10

Observation fe97cd66-6674-40cc-8114-c2d85b8c6e79 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Aligner: Efficient Alignment by Learning to Correct

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.357555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.357555Z digest=sha256:9c155584acfb092b24b008b9bb3b7c334b7bbd1a5e143d58179473954d13f07c

Observation ffc5257a-0570-4ec5-a019-8bb2b1a35488 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.273181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.361653Z digest=sha256:67110818dd2d1b37f6b6b028aeb28e2aca55cc4dacaab87890c5211a9ef25b0f

Observation 0c3ca86f-84c8-49a4-a559-a0d2e92c2df0 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.261362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.365601Z digest=sha256:37e53a77caed0ed98c92d652e1739a8c8165eef408cfdabe6c178337b58ad2dd

Observation 1cbd7cbe-0eea-4de5-b983-a42d7b8e4fea · outbound

This paper cites and Langford, J.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Langford, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.249699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.369026Z digest=sha256:182f6fcba2f0308a832431017fa63771c248f3f21558b91154f9a3cb55cfd9cb

Observation ef741c55-a3d1-444e-949e-cf7bdb033418 · outbound

This paper cites C., Jain, R., and Nuzzo, P.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation C., Jain, R., and Nuzzo, P

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.238206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.372814Z digest=sha256:9000ccf99fefee92a2a2e32ff8c2c9e25cbcb770058ee52c55634ba2b4eccecf

Observation 305c086e-6476-4c01-a751-90a7ed9ea8bc · outbound

This paper cites Reparameterization gradient for non-differentiable models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Reparameterization gradient for non-differentiable models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.225583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.376100Z digest=sha256:78b1d17013330f73a59a731a3ed04597021dc82f14bb92374be6c30ee31713de

Observation 376d8cbb-5fde-4d28-bcf4-32f3a7da07f5 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained variational policy optimization for safe reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.379471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.379471Z digest=sha256:8095682aeed72b288944fb5385ce3c7ce429857bec3ffebda24bd395f2e8d10f

Observation c807d510-2679-498a-a0dd-ba2240b1b7dd · outbound

This paper cites An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.383211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.383211Z digest=sha256:cd2720d1662da9bf287fc8310760db9f3841b8d46ce9fe6bc9371855036289db

Observation cbd0547a-70e9-47d5-906c-7bfe07d721c8 · outbound

This paper cites Gradients are Not All You Need.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Gradients are Not All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.386945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.386945Z digest=sha256:38ad847a10946b934021c1bee0ada8ffb4c2040509baf019f88b4cd599a86421

Observation 2bc24076-c5a0-407f-b5b7-b541570fd9c8 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Monte carlo gradient estimation in machine learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.204095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.390858Z digest=sha256:b9efbd68d2b9c93e1307fd21b10fc1a0fffcd7dc410b9d8c9f301488f8b76df4

Observation 49bd8828-605e-4244-a788-311f55409221 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.191282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.394340Z digest=sha256:7e530701ec99397677b86cc09cf61ec7ff46b24ee2e6d56e7a628880fa2f6d1b

Observation 7acdd6df-4a9a-4ef2-bff5-8f5a322e9b8b · outbound

This paper cites A focused backpropagation algorithm for temporal pattern recognition.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A focused backpropagation algorithm for temporal pattern recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.178946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.398561Z digest=sha256:df8890ec078e082adc117b4c54408ad7b1e50594e7f925db70f5c3d2f64f6713

Observation a63348ae-db32-4cce-ab11-1df098d90384 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.165377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.403481Z digest=sha256:3ca3a4381db0a14d449d3877f32a961abbc2b5efee6cab021154598f9fd8d95b

Observation cbbc98fb-67aa-411f-bb0a-9dd7e69882c8 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.151741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.407548Z digest=sha256:d57c31b6d96166193447199ccef8c2d57debcf3d76f196a3584f39eb509241c3

Observation bf742935-2efa-4048-97d0-0bba24bf954d · outbound

This paper cites Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation

Reference 37

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.584281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.412556Z digest=sha256:c78357ca6257ea57f586f55bfe2a16fdfffd6b4019cc39c3dacf4a88e335b7b7

Observation c1761bfb-2b04-466e-9cb2-61bcc64c4769 · outbound

This paper cites M., Smaby, N., and Cutkosky, M.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation M., Smaby, N., and Cutkosky, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.134741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.418519Z digest=sha256:2c853e8420918567e9a9928c594689538e4d4d29a65a998ec656384e6d5f464c

Observation 712e9719-cd9b-4042-80e5-8bb490350a44 · outbound

This paper cites Training language models to follow instructions with human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Training language models to follow instructions with human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.423648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.423648Z digest=sha256:81c5f4910a18087bfece4b416fb6f2dff0a6f9f4433ec56e75f90e27cee31dfd

Observation 139bfd47-8804-48ff-9cb5-13faf20c5b9b · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.432253Z digest=sha256:e3ced99f30eb1424c3f2224a496dfcefeb9a0ef627b0fc057ba9435500ea4a8b

Observation 3114b496-2514-448b-852b-9f77beb9670a · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.095465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.437224Z digest=sha256:51fa8eed3794b4deb7ff5d4e4f9168e66e03891f547aef5d4f0fa678726c1b59

Observation 633aab94-066a-433e-8a49-dcfaa483f501 · outbound

This paper cites and Barr, A.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Barr, A

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.082293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.441781Z digest=sha256:c70ecd78013475b541025f2da97af34734a2215dc932f73055313ccf179d475c

Observation 682a5e1f-371a-4840-b23b-9977cfb85249 · outbound

This paper cites E., Perescu-Popescu, L., and Mastorakis, N.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation E., Perescu-Popescu, L., and Mastorakis, N

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.067430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.446967Z digest=sha256:9c46d94c4ecb0b782cdc62a268d4809a1d10a67ce0d3df7640280f418c0e3fed

Observation efaa3de5-6dd0-45e0-9e11-2630796a9d4a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.451697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.451697Z digest=sha256:e686a1c865be7efdd92b038c8faceb2475a5c1f59fe35dabe1d8213005b6c100

Observation d5a0be68-f149-4d2b-a9ca-6fb830cd6854 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.456035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.456035Z digest=sha256:409171f5aacf6e57843f81f19751ca27b5da455afb0effe3991afa499d3cc105

Observation 5af2260d-6e97-4eb8-8543-0fd2555934bb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.460994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.460994Z digest=sha256:1d7587172d306e2d3acfe8d2d52b332e56271e9d954a725cf5322141ed10e4a3

Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.465505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.465505Z digest=sha256:e1e6c3bccc9bd73e06fb21630e38fb50370ef117ecd38ed9491eef668504cbd3

Observation 3a617aa4-eb65-40b0-9b18-09692828d39e · outbound

This paper cites Trust region policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trust region policy optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.470091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.470091Z digest=sha256:bc767e864ff1c23871db767175a2e66ccb8861f75eb1be63dbaf6b30fc7fcdfb

Observation 30eb5d28-521f-42c7-9a2d-90b69df97748 · outbound

This paper cites TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.676622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.473886Z digest=sha256:80a0de5d33d80ade324dcd6a55ff5ef9756937a83b7d047f152474daaa5e1b61

Observation 2164f7e4-e74d-4d9f-9cc3-e67046c62bb1 · outbound

This paper cites J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.478429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.478429Z digest=sha256:3ac452a94491f3c2eaff7066e34ff8a109458085038a68810fcb0caee2057176

Observation 5d2a9931-94a1-4178-8e73-2206990cc789 · outbound

This paper cites Mastering the game of go without human knowledge.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Mastering the game of go without human knowledge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.482832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.482832Z digest=sha256:2d9ce8603e4d0086b46d5e814ec45e155e51c1af25dd4940c4479b5f79cc1144

Observation e4475aa1-6272-45d6-8567-4328e096c74f · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.008450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.486674Z digest=sha256:9e7f408ee6c18e0b8d5b505a0c24b4a992ace8ce42dc5f492723b9b8f7dbe04c

Observation 3b87282a-bebf-4bd3-99f1-0c0868e88f05 · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Responsive safety in reinforcement learning by pid lagrangian methods

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.994990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.490259Z digest=sha256:9bb6eda6fcbc7abd3d4c53c01b3154a087bd41fea06c624805f86fc81acc6103

Observation 70b62e8d-cedb-4af6-9ad0-e4ab9cc69157 · outbound

This paper cites J., Simchowitz, M., Zhang, K., and Tedrake, R.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Simchowitz, M., Zhang, K., and Tedrake, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.980047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.494887Z digest=sha256:b1f5a77bdbe55796279c40fc224ed4a05ba52193db2c3da116698c7f352c19f8

Observation b60e1a78-6ea9-4152-85f0-a6a952922a4b · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.499664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.499664Z digest=sha256:7d51029cf43a38715d3458fed3440b6175aeac6f39e5e2d5ea5c9381466fa5e5

Observation f5a7f6ff-f92e-4263-875e-631c4c3bad3c · outbound

This paper cites W., Wang, T., Shang, Y., and Wu, Z.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation W., Wang, T., Shang, Y., and Wu, Z

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.956471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.505091Z digest=sha256:9c461374fc1fdbdeffa0bf035ef6711897bb6cd54b6056cf28826e95561fd565

Observation a320c451-bbe7-4f3a-b23b-6681bab8ddc3 · outbound

This paper cites Development of a humanoid robot control system based on ar-bci and slam navigation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Development of a humanoid robot control system based on ar-bci and slam navigation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.943923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.509463Z digest=sha256:d1980140a56c0ce14b17df1e2fa0b87057d6963e3cf4bdd12174305d7674772f

Observation cd42231a-daa8-4f4a-9662-c64b4836664a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:56.931716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.515075Z digest=sha256:ee57331d0fc0f8bf548bc261ea6b2e4c7ae335d2b733cb5c334977c15fbf4545

Observation 39328285-6824-4b56-97bc-6b1ea74a029b · outbound

This paper cites FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.521513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.521513Z digest=sha256:0ae5ae05c5ac073733c6a2f8e290b0e34dfadf63ad9f592daddc4443235e3533

Observation 9f83f98c-d743-4be3-8402-224876a928a8 · outbound

This paper cites Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.526507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.526507Z digest=sha256:bb0ce98d863c1945791bc0284b8469600c9c8155d05a36908a2a915811d497fc

Observation 1df34a0f-ee63-4136-a3d1-302b68d056d5 · outbound

This paper cites A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.632413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.531090Z digest=sha256:da08569087273a5dfd80462840289d7a46ea84f1915b1785dfbe5940f179db53

Observation 73e89b4b-0b3d-413f-a788-dfa37e890dae · outbound

This paper cites Constrained update projection approach to safe policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained update projection approach to safe policy optimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.917830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.536115Z digest=sha256:83ed7de83a32ed9c69f711706fa25e2d19b08f586ad6219f9fd0de85e20a2c40

Observation f7d4c6dd-6a6a-4971-9ab3-bf5fcb76046f · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Projection-Based Constrained Policy Optimization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.540973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.540973Z digest=sha256:7fcbced529d583a398db6569b82cdc45b909348f3081b37ae619bec5e40728ca

Observation 4b6b7e47-0add-45be-89ac-407b5fee5bee · outbound

This paper cites First order constrained optimization in policy space.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation First order constrained optimization in policy space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.901732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.545141Z digest=sha256:653321e144c9802becb65e064a5eacb124ccc5320a43c21e48c7264da6564e93

Pith citing papers

No inbound Pith citation observations are available.