Pith. sign in

Paper Citation Record · LEDGER

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2303.15810.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.15810 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T23:47:45.866615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.940883Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 369249ae-67b4-40de-85d7-24c8df8067e2 · inbound

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies cites this paper.

IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 50

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T13:48:36.611406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T13:48:36.369334Z digest=sha256:34ae0266548d64b1d265a52d8f2d7079fbd4c291dc14b9c073b57d9540ed766b

Observation 0c67fefc-1819-42ba-9415-5af47f1e6ea1 · inbound

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning cites this paper.

BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:52:15.285443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T10:48:28.980868Z digest=sha256:e72fd3d3e58d9b1e50babf27ca0175e1ff90d90212a79c3047233f317f2bacfc

Observation c79745c4-080c-4a8b-9206-447d4de34d79 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:9005b39c8cde0d01a7feb215bc7e064d8f50916b7b5312f41dbed30065e1ad84

Observation 93f16d80-24d1-46ed-b157-e537401a1850 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.622646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:cbcbada97932f9f6850d42a7285d0bc65cae959f94cbc97f0ca3b7030219fb0f

Observation 9b03b8b9-acfa-4d1a-92c4-483729c5c34a · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.618145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:5b888bbd08604ddb7d4abca313e14f2dd697b0a993ecf5e81a5356dc13d8c53c

Observation 48544bef-20c6-44af-af27-d789c3d80a4b · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:56:00.631554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:2e4f41efece4df9a82aeddfd323566473a7ef62a1df625b30a8c60fe1fb3f49e

Observation 38d69b95-3dee-4ae3-ae56-96d2a6576225 · inbound

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning cites this paper.

Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:25:45.840469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T21:45:43.298829Z digest=sha256:da495b4d48a66a9d58c1899f6489d3fadb7b59ed7e41ca76ff8a1e9c0ea27493

Observation 8b6381b3-ad22-4a49-82f8-610a1812c1e6 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.990694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:ea068120334503884ef70e0de7cc48371e1714f6b88eecf219152acaac73542d

Observation 1cf9ae9f-5987-4cd4-9237-5260c287e35d · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.942518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:e8e2ec0e7cd11d309ed2804f032c731e6b7b4933ba5276491200c18d4fc437d4