Pith. sign in

Paper Citation Record · LEDGER

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2507.22640.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22640 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:23.413516Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47d0afb2-e201-4c9c-91b8-555d6cd64d96 · outbound

This paper cites Reinforcement Learning: An Introduction,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Reinforcement Learning: An Introduction,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:28.983136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:16.903656Z digest=sha256:38a55a0929add889154829ef6c8de827bc8f8292ee50d3d83972c849f1518698

Observation 774554ce-b087-4f03-8bc8-5edd85b24f8c · outbound

This paper cites From automated to autonomous process operations,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction From automated to autonomous process operations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:28.773197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:17.019529Z digest=sha256:78e7866ee9ee7c1eb22ff03dd14205d6a5ea4f38f10990654dffbfc43cce5dff

Observation 1c1f6956-1bd8-4db8-b9ec-41dfddbcc373 · outbound

This paper cites Concrete Problems in AI Safety.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Concrete Problems in AI Safety

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:17.173306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:17.173306Z digest=sha256:c780a56343bb05aaf3757de0cb03b5c7987da4cd0470a2aa08bfb5b4c09266d9

Observation 508c8ec1-2f66-422f-a823-a991a2c02b1b · outbound

This paper cites Optimal grade transition for polyethylene reactors via NCO tracking,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Optimal grade transition for polyethylene reactors via NCO tracking,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:28.524681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:17.315159Z digest=sha256:a2337a710dabc622a9161b59412f8dabfcd8fcf17a2ae81604cac66b5d1a19ba

Observation c5c74b45-b2e3-4584-8979-918730fc3169 · outbound

This paper cites Iterative learning control-based batch process control technique for integrated control of end product properties and transient profiles of process variables,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Iterative learning control-based batch process control technique for integrated control of end product properties and transient profiles of process variables,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:28.351490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:17.502796Z digest=sha256:24d5116d2bc95cb15adf6c04280e00930f18318dc1b7a3fbd54f1eff2ab8b813

Observation ec392a86-50b6-4a89-922f-37f1aa9bb745 · outbound

This paper cites Integrated scheduling and dynamic optimization of grade transitions for a continuous polymerization reactor,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Integrated scheduling and dynamic optimization of grade transitions for a continuous polymerization reactor,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:28.139301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:17.711036Z digest=sha256:8422b55fb8e63c47ef85731b9bf80e32936fd33df30e41567d999e38f5739cc2

Observation 99ea38c8-7d60-44dc-b15a-e78182ec8bf7 · outbound

This paper cites The general problem of the stability of motion,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction The general problem of the stability of motion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:27.895469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:17.887751Z digest=sha256:194d2386ff081be2f33b6c61f39e7c8743d46d9afdde4c24633c642df0d77005

Observation 37dbd767-c63b-4c32-a7f4-cfbedc629599 · outbound

This paper cites an unresolved cited work.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T11:34:27.725152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:18.102389Z digest=sha256:17ac49c7849ff916d0468bae48db17fd35ee8c2b87839d4f363267681bd41504

Observation eb5cd566-8ea5-4177-b7a8-d22666238a99 · outbound

This paper cites Input Convex Neural Networks.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Input Convex Neural Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:18.321318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:18.321318Z digest=sha256:958593960a3ccc36a9f84fb1c5618847a4e2428440279c27c34129fc47a0e7b1

Observation aeeaa0b7-fca2-4bff-938c-093e73a7bd10 · outbound

This paper cites Safe Model-based Reinforcement Learning with Stability Guarantees.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Safe Model-based Reinforcement Learning with Stability Guarantees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:18.448416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:18.448416Z digest=sha256:8d3bcaa8a47fe8aaef39f1beeea8f7071942b5dab1022b6562949025f1722140

Observation dc70046b-ce11-4f08-afec-de8e3b61fd2f · outbound

This paper cites Control Barrier Function Based Quadratic Pro- grams for Safety Critical Systems,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Control Barrier Function Based Quadratic Pro- grams for Safety Critical Systems,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:27.486926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:18.607406Z digest=sha256:1c194b089bbd75911d0250e4a371f5bb8a815045844861b5ebc786abfbd04a33

Observation a584b792-8643-44bb-8e73-9c5b521064e9 · outbound

This paper cites Safe and Stable RL (S2RL) Driving Policies Using Control Barrier and Control Lyapunov Functions,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Safe and Stable RL (S2RL) Driving Policies Using Control Barrier and Control Lyapunov Functions,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:27.279331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:18.749875Z digest=sha256:15c86a99d0d00cb0018251cddb90680366971c2a9dd4b9ee962fe92b7169675f

Observation 6fbb22bf-b6ad-4ad6-a4a7-bd1ad960f3de · outbound

This paper cites Constrained Policy Optimization.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Constrained Policy Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:18.893392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:18.893392Z digest=sha256:3f02cd67f9525d1ea687ef9ef06c54035a12d5dc7cf7e3d2c4e68f8e6d274bf3

Observation 3d6583d2-0b5b-4715-98a1-3576e07c9e16 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Safe Exploration in Continuous Action Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.104125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.104125Z digest=sha256:12518ea646c17d6e569f7c801a5a3c1297da3b95894b75b35fb38a94aa8c55a2

Observation 10db63f3-0676-4291-8c36-01bea85205dc · outbound

This paper cites Conservative Q-Learning for Offline Reinforcement Learning.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Conservative Q-Learning for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.197355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.197355Z digest=sha256:23d2ca220cc7a7866f65c8211bbbecc7aac5d2d7e61debc83463ac49803c642c

Observation c007703d-cd76-4faa-97ee-60b7af00f21f · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Offline Reinforcement Learning with Implicit Q-Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.396964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.396964Z digest=sha256:da16e027f5965ee37c7bf31a738e6c71a2f4607645d9ccca933b2527b81e2267

Observation cf93d61e-21ef-4c72-b842-e954f85b855f · outbound

This paper cites MOPO: Model-based Offline Policy Optimization.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction MOPO: Model-based Offline Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.560544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.560544Z digest=sha256:9247aa4977724c188450e06686ab3a90d8686a747d7c7f9e29bda3c3d9396741

Observation 4ccfd0f9-726d-4af3-afb6-4699c5e56739 · outbound

This paper cites Actor–Critic Physics-Informed Neural Lyapunov Con- trol,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Actor–Critic Physics-Informed Neural Lyapunov Con- trol,

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T11:34:24.277113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:19.741550Z digest=sha256:53be54ad98c1023410dd571787f89644fd6c58b407d230a77877b789a05ed282

Observation 96778281-0ef6-4a6c-8dc0-c64a2fd2973b · outbound

This paper cites Distributional Reinforcement Learning with Quantile Regression.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Distributional Reinforcement Learning with Quantile Regression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:19.925848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:19.925848Z digest=sha256:c41ff574998dd60177504bbc46f907e90feebac64e2ef5af06ef871f40ed955d

Observation 370834ac-9400-41f1-a108-4603f8744f26 · outbound

This paper cites EKG-AC: A New Paradigm for Process Indus- trial Optimization Based on Offline Reinforcement Learning With Expert Knowledge Guidance,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction EKG-AC: A New Paradigm for Process Indus- trial Optimization Based on Offline Reinforcement Learning With Expert Knowledge Guidance,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:27.022358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:20.077489Z digest=sha256:38282f4e2dc0a5ee09ca970af82e4604db6f40c0dccdad37ad58e924cc024758

Observation 3826ad44-7d6d-4395-9796-201a044176c5 · outbound

This paper cites Optimal Control Via Neural Networks: A Convex Approach.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Optimal Control Via Neural Networks: A Convex Approach

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:20.306372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:20.306372Z digest=sha256:5200d2189e21a181e461f9951cd1ca8cadcf5027ef51359b61412bf194fdbe8c

Observation feab1efd-3c40-4dfb-9e6a-13241dd895b7 · outbound

This paper cites Differentiable Convex Optimization Layers.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Differentiable Convex Optimization Layers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:20.442610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:20.442610Z digest=sha256:e01ead5892cca3f282bbe51f072977b3f8286963b3be771a2d0db88dde91922e

Observation cb8b85ae-3c1d-4330-9859-32c9be6729d1 · outbound

This paper cites OptNet: Differentiable Optimization as a Layer in Neural Networks,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction OptNet: Differentiable Optimization as a Layer in Neural Networks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:26.677000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:20.590273Z digest=sha256:3e1f78ace5068c5f6635ed9868b7b067d1c6ac3e71f767e4d94285c0188efcc7

Observation 8d8fd4b8-b759-4fe9-89ab-822f849d849d · outbound

This paper cites Polymer grade transition control using advanced real-time optimization software,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Polymer grade transition control using advanced real-time optimization software,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:26.431238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:20.937196Z digest=sha256:15e497f8181e71cae491a52fcf8d2f8e5a660786addc34a16c902fcf5ec5eb48

Observation f69c0586-8e19-4c5e-bd61-6aebe21db289 · outbound

This paper cites Polymer grade transition control via reinforcement learning trained with a physically consistent memory sequence-to-sequence digital twin,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Polymer grade transition control via reinforcement learning trained with a physically consistent memory sequence-to-sequence digital twin,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:26.193384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:21.154028Z digest=sha256:9d5347ecba355d4cfe484ca542fcdcd1e4bf855fe6e75bd2998bf4e5ee98ca0b

Observation 0ba9b1af-a386-4ca0-b2af-30e290e13a68 · outbound

This paper cites A benchmark environment motivated by industrial control problems,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction A benchmark environment motivated by industrial control problems,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:25.938005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:21.296216Z digest=sha256:e2be0bc45db6292d1eb2c971fc587e9299061e004019ee2a575715ce7a9d16cd

Observation 9e3f99c7-1864-46b0-97c3-b30cd9b1c067 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:21.442854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:21.442854Z digest=sha256:e39c28116347185461fb40a9c4b75f5187b8d3537c3d91b43e5204289252b645

Observation ef465ad6-a07f-4c36-8315-f2958d7f4a43 · outbound

This paper cites PC-Gym: Benchmark Environments For Process Control Problems.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction PC-Gym: Benchmark Environments For Process Control Problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:21.584194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:21.584194Z digest=sha256:be20daf48e48513da8e400b7717d0dd8c4f9d432b9317bacad6c2b2ead5303e5

Observation eb487b2f-da8a-49e3-86a9-7867a95ee095 · outbound

This paper cites End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:23.746247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:21.799762Z digest=sha256:e1634e693806f8b8f22113db9ed5c5bad6ad0aaa7ae88f853e0a7afe05d38d76

Observation 203fd197-31c7-47a0-99e1-e90cc81eff2c · outbound

This paper cites Offline reinforcement learning methods for real-world problems,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Offline reinforcement learning methods for real-world problems,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:25.709828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:22.030643Z digest=sha256:d6f58c40e87681e11f59e8b772960ab67d6acfad30c5619f4d32eeb763f26918

Observation dae50e37-7c3d-4137-9e10-976a79f04421 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:22.256553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:22.256553Z digest=sha256:c908008dff40de8f1df11bfb8ec257153df04e1fd23625019150e2ecea22a74e

Observation fe6784ac-7e2f-4dcc-b52e-8fa68cd782bf · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction A survey on offline reinforcement learning: Taxonomy, review, and open problems,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:22.483538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:22.483538Z digest=sha256:c4be44ed10585688be2d84532bf3c1a9b8cf12fb98d65f3ad266fd70795547f6

Observation dce9edd8-9547-4c5a-8e76-9faad7a9eafd · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Stabilizing off-policy q-learning via bootstrapping error reduction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:25.474108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:22.678861Z digest=sha256:8502b844b7e30ce3f321daaec6553a7c4f85a31660ff9fae42fc2c61befb5423

Observation bf7f5e39-0013-46cb-9550-a3bcbe6650ab · outbound

This paper cites Deep Reinforcement Learning with Double Q-learning.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Deep Reinforcement Learning with Double Q-learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:22.841058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:22.841058Z digest=sha256:bb3f8ec2667385f6f8084c14acfcbb707eb6b9ea720d2517488644925ee5d0fc

Observation f22b21aa-dcd3-4daa-a3f5-ab7a82bb7621 · outbound

This paper cites Human-level control through deep reinforcement learning,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Human-level control through deep reinforcement learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:25.347910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:23.006757Z digest=sha256:56c8ef31b9b173d412ab288d8c8f5eca35de7e48e2a186a82ad0fa2536f77f08

Observation 2d55ea92-d8cd-437c-b857-b4e5247f3f56 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Behavior Regularized Offline Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:23.213288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:23.213288Z digest=sha256:6989f06d92c1c1cb193f265cff1d89973dd86981cdaffef58e74bf3f113bf129

Observation 176e6c8e-b993-44a6-a594-cc0c91212bb0 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:23.313489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:23.313489Z digest=sha256:f5327801ea8e90c55124ff83e3fdafc5e086be88cab2f2167a2783c8937b6bfd

Observation 09431ac7-47f5-49c0-a3bc-9ec13e86c089 · outbound

This paper cites Comparative Study of Machine Learning and System Identification for Process Systems Engineering Dynamics,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Comparative Study of Machine Learning and System Identification for Process Systems Engineering Dynamics,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:25.000267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:23.365230Z digest=sha256:ce3c705638bd9c3ac6ab54a8ab27bb0b0e9e16959f8d22586b56de58d31c1719

Observation bc6fd994-6286-4aed-8703-3ba5690351f0 · outbound

This paper cites Polymerization reactor control using autoregressive-plus Volterra- based MPC,.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction Polymerization reactor control using autoregressive-plus Volterra- based MPC,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:24.645311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:34:23.413516Z digest=sha256:c56dad5417bab9c31677b7d6cda927303063b4d4e3d455d37262bdbd9a8ef3e5

Observation 8ff2815d-9103-4710-be2f-c89055b741cc · outbound

This paper cites OptNet: Differentiable Optimization as a Layer in Neural Networks.

Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction OptNet: Differentiable Optimization as a Layer in Neural Networks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:20.749099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:20.749099Z digest=sha256:bf2e8bb3d86aaac25ffb86ef4ff6816b96a07e423ebe79227c819a6641fe4df1

Pith citing papers

No inbound Pith citation observations are available.