Pith. sign in

Paper Citation Record · LEDGER

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2303.15810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.15810 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:17:38.668063Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.940883Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 369249ae-67b4-40de-85d7-24c8df8067e2 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T13:48:36.611406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:5d7dec99c2061155ff0b461f8ca6073e879d4375bd0f7c812948aa761f73248f

Observation 2205cb6a-9e03-4c92-abf3-b5123bf3f647 · inbound

Are Expressive Models Truly Necessary for Offline RL? cites this paper.

Are Expressive Models Truly Necessary for Offline RL? Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:40.570825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:40.570825Z digest=sha256:aa85de5f916e055aba22190b849f5aa3b6d845826c58a74502fd8f002c9fc8d0

Observation 20d6d736-a044-495b-8099-7d0881f00018 · inbound

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL cites this paper.

Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:30:49.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:30:49.142024Z digest=sha256:c0ae6230f6b05ae9d0a187022c580eb250932ebac52e7a69de35ed18dc8bd265

Observation aa0256fe-b7f8-435e-a6c3-a3eeb64e433b · inbound

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning cites this paper.

An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:17:38.668063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:17:38.668063Z digest=sha256:71b2221f5d84e2136c8c6806e742a500e4011bfe9153002d2002ab1b0ad623d0

Observation 5795690e-4847-4996-ba18-6a6adf07d008 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.427182Z digest=sha256:033787afef90cbf5014e8b443c1f4623bbc515a2da0412532d50b8f45737ebff

Observation 0c67fefc-1819-42ba-9415-5af47f1e6ea1 · inbound

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning cites this paper.

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:52:15.285443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T10:48:28.980868Z digest=sha256:adceb6e8d0c8f66415472ba6fa5f1e05c9461b64160ec959204685bb546f9f74

Observation a73e2a5f-1b28-4238-8d50-5f4e3eb823c2 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:27.299725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:27.299725Z digest=sha256:22e855b13770c2f156911cc1cf8bfcc42eba0275c4ad984090520964ad5bf894

Observation c79745c4-080c-4a8b-9206-447d4de34d79 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:856b575038851fc64b9a08554eea4417c80cfa357f6fdaf51516a37b4072cb45

Observation 93f16d80-24d1-46ed-b157-e537401a1850 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.622646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:c27cc686aa8f11b5dfed6fbc60aaac0557582fbf1190f73d6f5eefb489501ccb

Observation 9b03b8b9-acfa-4d1a-92c4-483729c5c34a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.618145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:17aab6b209eb2d60dd3bfcd9c1a81f46ae0eea860ac9ac2c8dccfdc05312b699

Observation 48544bef-20c6-44af-af27-d789c3d80a4b · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.631554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:14876b1bfb48ad498f4694953941f415ca0a67f333b00b41d563e1b765eb8428

Observation 38d69b95-3dee-4ae3-ae56-96d2a6576225 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:45.840469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:354ae24c4ba5428b7aa758db2a59fe962f39061dceebc209d15f06415598c5f5

Observation 8b6381b3-ad22-4a49-82f8-610a1812c1e6 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.990694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:2e35c38c9d4683802b4040b3bf2fb22700798f568c0d4a5ceb49d943b8ad3b34

Observation 1cf9ae9f-5987-4cd4-9237-5260c287e35d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.942518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:ae977b74016d6584b0452145cd57ae13a88c76c923195652503241c25acc5c4e

Observation 7f5d9b58-8389-43fe-a872-f437c8fd2cd9 · inbound

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL cites this paper.

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:30.253091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:57:30.253091Z digest=sha256:979a582f98c88d36fcfa57838a8ccdc5077931da91751adab3d801035bf64983