Pith. sign in

Paper Citation Record · LEDGER

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2412.08880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08880 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:33:37.721942Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 484da714-e99a-4520-b896-f63ea4633c16 · outbound

This paper cites Constrained policy optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained policy optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.528192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.528192Z digest=sha256:14f85109835a0e6258b25116882a230811db7e9137b38de8f3b545b77c38749b

Observation 614de274-5e77-4761-80bf-c28644564dc3 · outbound

This paper cites Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning in continuous state spaces with path integrals

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.307332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.534971Z digest=sha256:3e2a531337be6ae26b2e9171201c037ae495f7fa0faf9987b65d9e238d356cf2

Observation a791a083-db0d-40e6-93c3-9d19b8d87889 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.289432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.540639Z digest=sha256:98c061b6e8f86bb340f5fcbc9bb0b7343603ec5e354269d6a1d5aa9d4dc08f37

Observation 1db5df65-eda1-4c92-9873-86c4032d64d4 · outbound

This paper cites Constrained Markov decision processes.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained Markov decision processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.547006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.547006Z digest=sha256:e8fb6d8a624c48ca39221b43979d2ab369540544c9754fe8d11cfa6cf5847e54

Observation 636cafb2-0f7b-4233-ad75-f1032fce1913 · outbound

This paper cites Convex optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.552930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.552930Z digest=sha256:860ee451de0338306e1a6f6f639d4f161383063e0e805c01bc2c6a7f796274d0

Observation 93c8ef20-0422-4cfb-86d2-28ba9d322c2a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.558683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.558683Z digest=sha256:ccfce9e5769a6c3ad9498c727b6506a56bd96077d91bc703ce2607ab14776582

Observation 6b258255-db0c-43f1-a7bd-bd18cb70c4f6 · outbound

This paper cites Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Pybullet, a python module for physics simulation for games, robotics and machine learning, 2016

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.564869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.564869Z digest=sha256:5f55b8efc8eac029e36cf79ddc87729dcb945d562de92da2b4daf208d0e7c2a8

Observation 9256cbd1-13c6-4efa-8f2c-b3a2e828c268 · outbound

This paper cites Directional differentiability of optimal solutions under slater's condition.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Directional differentiability of optimal solutions under slater's condition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.226032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.569920Z digest=sha256:1e2f314bf1a5e80c45fb160d8eda85d2cca3cc1b40cafe4281ec6514c311bc45

Observation 71a64338-d0d0-4560-ad92-c457ed052236 · outbound

This paper cites Parenting: Safe Reinforcement Learning from Human Input.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Parenting: Safe Reinforcement Learning from Human Input

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:38.005117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.574512Z digest=sha256:249b1e0b1f1b82b9816123e91149b92d47420af738ec1110007fc5906ee60231

Observation 0deffe96-4c7c-4b64-bbfd-34b362135362 · outbound

This paper cites A comprehensive survey on safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A comprehensive survey on safe reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.580345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.580345Z digest=sha256:049d21edf05368b38a83bd40c945076e0225ccda5c5a8bf8343e78264e750a4b

Observation 1d1f6f5a-aa67-4549-8be7-782a782f91a4 · outbound

This paper cites A primal-dual augmented lagrangian.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A primal-dual augmented lagrangian

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.199580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.585226Z digest=sha256:caa19f168411bd8a431c081e42f532b52581c4bb773d9c281cef8ea50c2049f5

Observation c38cf8b3-58f2-4f41-9428-ffadf8c3f167 · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.592440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.592440Z digest=sha256:d863dacb0fa7af9917a3c80adfd4085fcc75ffffbdd657c4373798106bf29fd6

Observation 8643a124-d913-4c15-9b1b-750249d9db3b · outbound

This paper cites A Review of Safe Reinforcement Learning: Methods, Theory and Applications.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning A Review of Safe Reinforcement Learning: Methods, Theory and Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.597399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.597399Z digest=sha256:2ea7a59859ba5d135f9fd8ef787488f7214cf1c1eafa4d07a2a4864ace98468f

Observation 5ee64d99-f8cb-4f9c-9fd2-da6482429ed3 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.602358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.602358Z digest=sha256:d634d08398e935ce24f00880ad86ed3aeeda6ce201a7dae9ae6370e73726d7ae

Observation 866f1926-a3d6-47a9-aa85-735a1e3990d0 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.606897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.606897Z digest=sha256:853b2d896cb7261dbc6d7f7edc4bb3b2f81c4256d1cf6dc3955f5b35210fe0ed

Observation a0e80708-02c8-4cf3-8516-cc6f0c593526 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.612195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.612195Z digest=sha256:640e61fc45eceef71f24c9a193bddf794c08df7e59a1ed23c289c5bdf6b582a0

Observation 6f5cde42-7188-4a64-b4b7-cf4fa14bf058 · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.152117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.617099Z digest=sha256:a4b922e30247dcb88e51e309918d637dea6a75e743234d1f0a74bd3c6ed7c33b

Observation 14972fe8-b442-4a3d-9485-393ac1488e1b · outbound

This paper cites COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.622251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.622251Z digest=sha256:66587a0c6ad7bf1851b51b817dc29e1acab628c8837eadddd60f3f610d1348e3

Observation f31ca234-8c65-48e2-ae24-d88c464fbdfd · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.628179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.628179Z digest=sha256:00206c44f04c33a1c025871955e228c3086529a82db3b0995811eeb065293808

Observation 33bf6911-2cbc-4a33-befc-e5bf0747f0d2 · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.634783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.634783Z digest=sha256:c6ac232639f4ba316c265e445adc39282e42747ee518c32c71f26b302fc48bdb

Observation 1f0345e2-16b3-4acb-9857-da80b41b1e4f · outbound

This paper cites Constrained decision transformer for offline safe reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constrained decision transformer for offline safe reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.137434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.640165Z digest=sha256:a0bbf88d9ddb877dce6142ee78992a2e7652b06f9e879530b701e0ff853999a2

Observation 16d8aabc-e554-40fc-a42a-dc56489c2632 · outbound

This paper cites Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Feasible Actor-Critic: Constrained Reinforcement Learning for Ensuring Statewise Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.645535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.645535Z digest=sha256:bf184c054a5fd0e1a0df176b5f6d078c02bc1e7c63fa80e0bd6494276f9d9f0f

Observation cffaf56a-76eb-496b-89f8-c54cafa1ff6c · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.650823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.650823Z digest=sha256:a155cc421aa91cd52582c80acb37adc0881785d228ef64c6f04719a351134e0a

Observation a21cab1a-23d3-4346-9d9e-152408a0540a · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.656127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.656127Z digest=sha256:5e9b086f1659c4126ebfc065f56cdd7f8fd67b395f691423e8f8bc425821dd07

Observation e4adc729-f884-4777-bc6c-968f6d854617 · outbound

This paper cites Trust Region Policy Optimization.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Trust Region Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.661644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.661644Z digest=sha256:32022572af1d1175ead5c9932eb3e2a98aaa9f9a42c14319b77bd26b59d77513

Observation cccddb50-a740-4495-9b42-15adee25b7c5 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Equivalence Between Policy Gradients and Soft Q-Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.667394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.667394Z digest=sha256:e3f3affc356f9026666ab2c6e98f3fee8301a8bb9c41bd9491e4f71e52b206f6

Observation 3e162c79-8e30-482d-a879-d855d690bc2a · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Responsive safety in reinforcement learning by pid lagrangian methods

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.120946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.672685Z digest=sha256:996c0f48981cabf42a4b9a594b1374ccd4495c29bb743dc856b09030501fd6f7

Observation 6216db71-c306-4fbf-8ee6-4d198cb6b999 · outbound

This paper cites Constraints penalized q-learning for safe offline reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Constraints penalized q-learning for safe offline reinforcement learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.101582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.678677Z digest=sha256:516851224a4c498faad2402ca967a95489b47ea5b65595c1c0260a0b61aa7d83

Observation d0a69486-bf64-44aa-aa04-b430122cc777 · outbound

This paper cites Primal-dual stochastic gradient method for convex programs with many functional constraints.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Primal-dual stochastic gradient method for convex programs with many functional constraints

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.085356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.684038Z digest=sha256:78ba1fc1ee4c3e500fda457d71f97025e7b42c796a6a6143c70b7ba8fa58cbae

Observation 9eabaaed-dbf7-44b9-86f4-8d62f22dc0ca · outbound

This paper cites OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:33:37.808510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.689199Z digest=sha256:8d15e539ad30ce4112f2d1efe9e27e4244dd7040116d17088463714e7a7a684e

Observation 0e665c63-23ba-468e-a65e-3daf84174d1d · outbound

This paper cites Reachability constrained reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Reachability constrained reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.067425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.695599Z digest=sha256:6b290433b14328318e0aa6094ca326a1e96d2464e2f228374737a4b4382661c8

Observation 8f5e93cf-c4d0-48de-afad-450c830f1720 · outbound

This paper cites Penalized Proximal Policy Optimization for Safe Reinforcement Learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Penalized Proximal Policy Optimization for Safe Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.700628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.700628Z digest=sha256:6c6f114afa5d0f62471033b22889ac4f2626daf5ae5a135fed77216d38cf9044

Observation b5486b46-e347-47f7-9837-461c07eaa0ff · outbound

This paper cites Evaluating model-free reinforcement learning toward safety-critical tasks.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Evaluating model-free reinforcement learning toward safety-critical tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.050910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.706021Z digest=sha256:991c5233fdbbc1a15ab0c8f62e9d19b035c1556130872951aa94c3e7ef6e2db4

Observation da699856-f6a8-4d7e-b7eb-eca65762b5af · outbound

This paper cites First order constrained optimization in policy space.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning First order constrained optimization in policy space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:33:38.034039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T17:33:37.711703Z digest=sha256:1054bb5eb2ff75c727dba4d2f232168338f4df70575e7fbe767581325fdb00dc

Observation b59b261b-3528-4132-90d9-c51b8f8f0880 · outbound

This paper cites Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.716533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.716533Z digest=sha256:8920aa42f5cab4955a949daa80b6a5e1edd3151a096dd98e77a12fd9c491202a

Observation 0dc13ffc-13f9-44e5-94dc-f1d0ab8f545c · outbound

This paper cites Maximum entropy inverse reinforcement learning.

FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:33:37.721942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:33:37.721942Z digest=sha256:73bc875cddb248b2f922a933256c88072bb8b1dd7ffe8092b60e573f242eb148

Pith citing papers

No inbound Pith citation observations are available.