Pith. sign in

Paper Citation Record · LEDGER

Score Regularized Policy Optimization through Diffusion Behavior

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.07297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.07297 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:38:32.067287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.357585Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d21a81c5-2540-4161-b310-5a0870610646 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.479870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:f4a2b722dd110692ddefcb4bfea704189a6689c183beb9e1442c7952310f76a7

Observation f7d77ab3-7008-4262-a7e8-eb0341d5621b · inbound

Offline Reinforcement Learning with Penalized Action Noise Injection cites this paper.

Offline Reinforcement Learning with Penalized Action Noise Injection Score Regularized Policy Optimization through Diffusion Behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:32.067287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:32.067287Z digest=sha256:cb2f6a791985caef6ac2baf836bfb20ebb9e4922d3c48d948d43ac446588f70d

Observation d6fb52d4-c22f-446b-b8bb-f4bf0807ccaa · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Score Regularized Policy Optimization through Diffusion Behavior

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:2769ebb2148cfcc9b156ccc8912acaa78df3e03ecb207882b27c7d9e6ad5bfff

Observation 1d192dac-cf14-4ec6-9a65-16360f99cb38 · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Score Regularized Policy Optimization through Diffusion Behavior

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.558176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:ca61036199ced485f0c1e31a88f04882cc9bf1344af3f0190288cfacb29f7e35

Observation f5beaba2-02ee-4f9c-a0dd-0ae027d9e460 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.773020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:c7035c6b3bf7636f4a83872c46214e70484195de9795408f766d3dc79b1a708b

Observation cb921262-cf60-4f3b-8af1-a231886ac316 · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:48.333368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:e51f909edb18a6467293ef0b13ffd5b9e1b17a22f700f0ad20edb8fb9b2ced42

Observation db9e2e38-2a12-4fde-a37c-0685532fee09 · inbound

Path-Coupled Bellman Flows for Distributional Reinforcement Learning cites this paper.

Path-Coupled Bellman Flows for Distributional Reinforcement Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:16.247181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T02:12:58.528130Z digest=sha256:bce656c6f61652e59131a64ecb2972ae6da76df2c0cc00b3ded523429f20ecfa

Observation 3dfc10f2-7f7f-4631-bedb-5c612630c352 · inbound

SPAR: Support-Preserving Action Rectification cites this paper.

SPAR: Support-Preserving Action Rectification Score Regularized Policy Optimization through Diffusion Behavior

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:43:30.999353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T14:34:39.094753Z digest=sha256:2e187f51c6c048234b9217f8ef0184ebd1f7669000ddf7e41d60a4c5a2be3b1b

Observation f1b200e0-2084-477a-9a5a-43de682006b0 · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Score Regularized Policy Optimization through Diffusion Behavior

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.397932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:405cd02b8be1dbe19257f1e764ff614b88927c92ca1cd6b7e93a501bc99f02b9

Observation dc086fae-516b-4571-ac7b-7ddde9870d3e · inbound

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning cites this paper.

Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-Learning Score Regularized Policy Optimization through Diffusion Behavior

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.359291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:57:31.180928Z digest=sha256:07eb72c35363bb134faaa547aa5db0b8a3a09a93f031efed804d9a57c9612183