Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:33:37.721942Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2412.08880.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:33:37.721942Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 484da714-e99a-4520-b896-f63ea4633c16 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained policy optimization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 614de274-5e77-4761-80bf-c28644564dc3 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a791a083-db0d-40e6-93c3-9d19b8d87889 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1db5df65-eda1-4c92-9873-86c4032d64d4 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained Markov decision processes
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636cafb2-0f7b-4233-ad75-f1032fce1913 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Convex optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c8ef20-0422-4cfb-86d2-28ba9d322c2a · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b258255-db0c-43f1-a7bd-bd18cb70c4f6 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9256cbd1-13c6-4efa-8f2c-b3a2e828c268 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Directional differentiability of optimal solutions under slater's condition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71a64338-d0d0-4560-ad92-c457ed052236 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Parenting: Safe Reinforcement Learning from Human Input
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0deffe96-4c7c-4b64-bbfd-34b362135362 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A comprehensive survey on safe reinforcement learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1f6f5a-aa67-4549-8be7-782a782f91a4 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A primal-dual augmented lagrangian
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c38cf8b3-58f2-4f41-9428-ffadf8c3f167 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8643a124-d913-4c15-9b1b-750249d9db3b · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A Review of Safe Reinforcement Learning: Methods, Theory and Applications
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee64d99-f8cb-4f9c-9fd2-da6482429ed3 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 866f1926-a3d6-47a9-aa85-735a1e3990d0 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e80708-02c8-4cf3-8516-cc6f0c593526 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Conservative q-learning for offline reinforcement learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5cde42-7188-4a64-b4b7-cf4fa14bf058 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Optidice: Offline policy optimization via stationary distribution correction estimation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 14972fe8-b442-4a3d-9485-393ac1488e1b · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31ca234-8c65-48e2-ae24-d88c464fbdfd · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33bf6911-2cbc-4a33-befc-e5bf0747f0d2 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0345e2-16b3-4acb-9857-da80b41b1e4f · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained decision transformer for offline safe reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 16d8aabc-e554-40fc-a42a-dc56489c2632 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cffaf56a-76eb-496b-89f8-c54cafa1ff6c · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a21cab1a-23d3-4346-9d9e-152408a0540a · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4adc729-f884-4777-bc6c-968f6d854617 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Trust Region Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cccddb50-a740-4495-9b42-15adee25b7c5 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e162c79-8e30-482d-a879-d855d690bc2a · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Responsive safety in reinforcement learning by pid lagrangian methods
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6216db71-c306-4fbf-8ee6-4d198cb6b999 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constraints penalized q-learning for safe offline reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0a69486-bf64-44aa-aa04-b430122cc777 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Primal-dual stochastic gradient method for convex programs with many functional constraints
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9eabaaed-dbf7-44b9-86f4-8d62f22dc0ca · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e665c63-23ba-468e-a65e-3daf84174d1d · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Reachability constrained reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f5e93cf-c4d0-48de-afad-450c830f1720 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Penalized Proximal Policy Optimization for Safe Reinforcement Learning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5486b46-e347-47f7-9837-461c07eaa0ff · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da699856-f6a8-4d7e-b7eb-eca65762b5af · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning First order constrained optimization in policy space
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b59b261b-3528-4132-90d9-c51b8f8f0880 · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc13ffc-13f9-44e5-94dc-f1d0ab8f545c · outbound
FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.