Pith. sign in

Paper Citation Record · LEDGER

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2303.15810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.15810 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:27.299725Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.940883Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 369249ae-67b4-40de-85d7-24c8df8067e2 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T13:48:36.611406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:ef8e9aa1df7473105be8e138a7a4b289cf5d93bd2bc5d891da03a1da569184fb

Observation 0c67fefc-1819-42ba-9415-5af47f1e6ea1 · inbound

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning cites this paper.

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:52:15.285443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T10:48:28.980868Z digest=sha256:224261ff9c1f22f35db0efcf043ebb4d34569229a58d31a09aaf666945efe04e

Observation a73e2a5f-1b28-4238-8d50-5f4e3eb823c2 · inbound

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood cites this paper.

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:27.299725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:22:27.299725Z digest=sha256:50fc9bb8cd8b153f72071c609d05b88668b4376f70acdf809e9c70fc71c8e60d

Observation c79745c4-080c-4a8b-9206-447d4de34d79 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:9005b39c8cde0d01a7feb215bc7e064d8f50916b7b5312f41dbed30065e1ad84

Observation 93f16d80-24d1-46ed-b157-e537401a1850 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.622646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:f8337e7793230549dcac0775fd606dff7f3c9610928581c1c79395177a4b4bc2

Observation 9b03b8b9-acfa-4d1a-92c4-483729c5c34a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.618145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:5f094b578d5b8b8fdcd8eaebb9063681f496eba600cfd0cb17adf5c3f41887fa

Observation 48544bef-20c6-44af-af27-d789c3d80a4b · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.631554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:07d420b962c173f85ef035d37929319f3cf3363179e35e3561e879e16383aba8

Observation 38d69b95-3dee-4ae3-ae56-96d2a6576225 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:45.840469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:503c5f5e4fe03564d8227a7d0f787fea03b3f869989f23340b7be879bfae33af

Observation 8b6381b3-ad22-4a49-82f8-610a1812c1e6 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.990694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:ab986f76a27e333ca57e8080dfaf258ac319d192d7902598b763e346b234665b

Observation 1cf9ae9f-5987-4cd4-9237-5260c287e35d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.942518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:264114f042070aae8b2f45fba04e1243469f1ce969dfeb916df2820bea7fd431