Pith. sign in

Paper Citation Record · LEDGER

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2412.08880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08880 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:33:37.721942Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 484da714-e99a-4520-b896-f63ea4633c16 · outbound

This paper cites Constrained policy optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.528192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.528192Z digest=sha256:dbac4f5de9b6729f250bcc461190495afbb243a2915457a77444f8f1f4954710

Observation 614de274-5e77-4761-80bf-c28644564dc3 · outbound

This paper cites Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.307332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.534971Z digest=sha256:24a272ca0881017a7e088ab91e07fbdfffed691517e236c21360970fc1526c85

Observation a791a083-db0d-40e6-93c3-9d19b8d87889 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.289432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.540639Z digest=sha256:9e24a086bbe5ec56aea63c11605f7e8b7ea6024a4f52c2584c3b005c8f15f927

Observation 1db5df65-eda1-4c92-9873-86c4032d64d4 · outbound

This paper cites Constrained Markov decision processes.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained Markov decision processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.547006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.547006Z digest=sha256:b05843a4d27cbf5c1c853f3c38a7c25985be680a78fa2928f53b20b06105710f

Observation 636cafb2-0f7b-4233-ad75-f1032fce1913 · outbound

This paper cites Convex optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.552930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.552930Z digest=sha256:95d3c02ba2a8c43e5c5e2966757b00a696cda5f505741148dd2c4371121dc115

Observation 93c8ef20-0422-4cfb-86d2-28ba9d322c2a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.558683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.558683Z digest=sha256:2ec5990f97197489f55890ab9d86d1b2a88509053cacca242f1f2da5a0465453

Observation 6b258255-db0c-43f1-a7bd-bd18cb70c4f6 · outbound

This paper cites Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.564869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.564869Z digest=sha256:5b604a76e7533183e79b34691864ef57b0904a0b9bb10c2d08a2cfb7eed9c7a4

Observation 9256cbd1-13c6-4efa-8f2c-b3a2e828c268 · outbound

This paper cites Directional differentiability of optimal solutions under slater's condition.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Directional differentiability of optimal solutions under slater's condition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.226032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.569920Z digest=sha256:606f3782669ac80424807135371f26a2104c8b5bcd8f0e6ec68c98c89d7f7880

Observation 71a64338-d0d0-4560-ad92-c457ed052236 · outbound

This paper cites Parenting: Safe Reinforcement Learning from Human Input.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Parenting: Safe Reinforcement Learning from Human Input

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:38.005117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.574512Z digest=sha256:9d53292f37ee4c3748b259aa7994f1200673a72d92d461ce2da3517a7dc2a3bb

Observation 0deffe96-4c7c-4b64-bbfd-34b362135362 · outbound

This paper cites A comprehensive survey on safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A comprehensive survey on safe reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.580345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.580345Z digest=sha256:5f86abcd6a902cdc06992c639435959b693ceeb8f89ebefe5e85f78202963f46

Observation 1d1f6f5a-aa67-4549-8be7-782a782f91a4 · outbound

This paper cites A primal-dual augmented lagrangian.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A primal-dual augmented lagrangian

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.199580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.585226Z digest=sha256:f699335d0f88a6d9a3f29cfe5811f1f4e52af3ce7b100024eaf28c46a8351723

Observation c38cf8b3-58f2-4f41-9428-ffadf8c3f167 · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.592440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.592440Z digest=sha256:b1eac3e2b80d68535aba6515ea1f9b7799029aad2efee34b0cf7a90019e428bd

Observation 8643a124-d913-4c15-9b1b-750249d9db3b · outbound

This paper cites A Review of Safe Reinforcement Learning: Methods, Theory and Applications.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A Review of Safe Reinforcement Learning: Methods, Theory and Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.597399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.597399Z digest=sha256:1a15ffc889004889bae712e8f9cb11ff78c7f1a56a830e338b76f5985c162058

Observation 5ee64d99-f8cb-4f9c-9fd2-da6482429ed3 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.602358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.602358Z digest=sha256:5c274bf25cec5fbb596f02e1464c9d7fc9e633640f1c585c99107ef834ed3f5b

Observation 866f1926-a3d6-47a9-aa85-735a1e3990d0 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.606897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.606897Z digest=sha256:0b31848eba93d671e772335356ddd41bcee0d9f7795c5786fd048b121ae3c1ee

Observation a0e80708-02c8-4cf3-8516-cc6f0c593526 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.612195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.612195Z digest=sha256:604d925d16da44d304213044bd1715b5479a813e1040ba7845055733ef72690d

Observation 6f5cde42-7188-4a64-b4b7-cf4fa14bf058 · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.152117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.617099Z digest=sha256:b463b411e1bb0cef1ff83043fe3d13163fe6e61d93a48ca6bf6365f70d864964

Observation 14972fe8-b442-4a3d-9485-393ac1488e1b · outbound

This paper cites COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.622251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.622251Z digest=sha256:d1a1f9e2b67641ede39f9e95d1230adb6f9872bd46839ec1727daf08b4d61b88

Observation f31ca234-8c65-48e2-ae24-d88c464fbdfd · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.628179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.628179Z digest=sha256:cca1612a445dde717a323b0f2dad12ca80e2612bcf83427dce74ba3bf4ffdb5d

Observation 33bf6911-2cbc-4a33-befc-e5bf0747f0d2 · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.634783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.634783Z digest=sha256:d06d6ca22709d26a2588856983824b7307df0e6760e2915900c9cf5788fe5b47

Observation 1f0345e2-16b3-4acb-9857-da80b41b1e4f · outbound

This paper cites Constrained decision transformer for offline safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained decision transformer for offline safe reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.137434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.640165Z digest=sha256:57df5f872127c5a04392c6b117b742636ea1f2d477ee2b16954835a5c7b2676a

Observation 16d8aabc-e554-40fc-a42a-dc56489c2632 · outbound

This paper cites Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.645535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.645535Z digest=sha256:9219ae74c5455844b0e4062317b3a738de1ce259443be1d6994db588801a9e0b

Observation cffaf56a-76eb-496b-89f8-c54cafa1ff6c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.650823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.650823Z digest=sha256:209b19e5817d309b608d8e823c8c25a20a56f5a7fde5851e599bf97a28be896e

Observation a21cab1a-23d3-4346-9d9e-152408a0540a · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.656127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.656127Z digest=sha256:bb8055d3fc9d7a330c8834685b1c5f8d59a8106331ade15bb5a8a179bb316baf

Observation e4adc729-f884-4777-bc6c-968f6d854617 · outbound

This paper cites Trust Region Policy Optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Trust Region Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.661644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.661644Z digest=sha256:b7874cd89e6875b8761a8b25874a32ba0b57276124d10c0d95a838efdacfcaea

Observation cccddb50-a740-4495-9b42-15adee25b7c5 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.667394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.667394Z digest=sha256:ee110df0bc784fb3017dbd13c3012a4c22137b2c21a9c350f4d7d396e89323dc

Observation 3e162c79-8e30-482d-a879-d855d690bc2a · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Responsive safety in reinforcement learning by pid lagrangian methods

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.120946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.672685Z digest=sha256:5fe9b24b3fe707a55bfca21b25d0b7d846ab9caf0b273dd5aafe5eb10792dc46

Observation 6216db71-c306-4fbf-8ee6-4d198cb6b999 · outbound

This paper cites Constraints penalized q-learning for safe offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constraints penalized q-learning for safe offline reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.101582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.678677Z digest=sha256:978dfc01cf34202bb3c8b9df66e2d7699c5e9e743831109b44d6e2f3d769dc2a

Observation d0a69486-bf64-44aa-aa04-b430122cc777 · outbound

This paper cites Primal-dual stochastic gradient method for convex programs with many functional constraints.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Primal-dual stochastic gradient method for convex programs with many functional constraints

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.085356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.684038Z digest=sha256:2ce12888e92b68efbc50c6d0b973793f630c9fc3bfcd1cfc4f8ac8ecd2c032e3

Observation 9eabaaed-dbf7-44b9-86f4-8d62f22dc0ca · outbound

This paper cites OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:37.808510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.689199Z digest=sha256:d1bacf4ce054d207cd34121a153e4964183f06edbd1404d9d9c4223af1df4bc8

Observation 0e665c63-23ba-468e-a65e-3daf84174d1d · outbound

This paper cites Reachability constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Reachability constrained reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.067425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.695599Z digest=sha256:d406198a1528ae7240ed03b117b5626a0672d5d42a78bcd30239720dd5f9b0be

Observation 8f5e93cf-c4d0-48de-afad-450c830f1720 · outbound

This paper cites Penalized Proximal Policy Optimization for Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Penalized Proximal Policy Optimization for Safe Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.700628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.700628Z digest=sha256:9b2ccec88afc5dcb667c7c00c3052d48897676324c38b0edda648c28a6e41130

Observation b5486b46-e347-47f7-9837-461c07eaa0ff · outbound

This paper cites Evaluating model-free reinforcement learning toward safety-critical tasks.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.050910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.706021Z digest=sha256:f77a3cf3d9085fc9cff057c1fce376ca67431bfcd384fb5883c8b8bed438f876

Observation da699856-f6a8-4d7e-b7eb-eca65762b5af · outbound

This paper cites First order constrained optimization in policy space.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning First order constrained optimization in policy space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.034039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.711703Z digest=sha256:6888f0bf9f8cd2d43b116cc22cb7fcc9accefe512454c9c936b268afe885cf59

Observation b59b261b-3528-4132-90d9-c51b8f8f0880 · outbound

This paper cites Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.716533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.716533Z digest=sha256:0904273ed22d233afa174043d81a7db4cde8e021cb9674b2431801efcdbcaef6

Observation 0dc13ffc-13f9-44e5-94dc-f1d0ab8f545c · outbound

This paper cites Maximum entropy inverse reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.721942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.721942Z digest=sha256:4bd35a5712c2da8d26a3ffee67037b2309e4887ecb3e1f4db673b16593adfa1d

Pith citing papers

No inbound Pith citation observations are available.