Pith. sign in

Paper Citation Record · LEDGER

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2312.11752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11752 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:00.346736Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.937596Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 16c179a2-3336-4052-94c8-4f91798adf1a · inbound

Diffusion Policy Policy Optimization cites this paper.

Diffusion Policy Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.975258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:444df2ce5387b7b0f95a47f5dddebdb29bfcf839480da3e58021145aabcfcaf5

Observation 07e6a49b-df04-46df-a177-0ec39454bb38 · inbound

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation cites this paper.

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:48:30.730504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:48:30.730504Z digest=sha256:511050e1893f32977460ebbf4dd245dacae7aaef68836cab64adf991c08ad773

Observation b399cbd0-d316-495e-9e4b-ca909ce2e3a7 · inbound

Generative Diffusion Modeling: A Practical Handbook cites this paper.

Generative Diffusion Modeling: A Practical Handbook Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:49:35.453168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:49:35.453168Z digest=sha256:da0152dd0927d92b84570b6ada34e2bc642a6e160805b55338c42220eaab836f

Observation f5b3d552-8eee-4bcf-b728-92f8aaadc188 · inbound

Efficient Online Reinforcement Learning for Diffusion Policy cites this paper.

Efficient Online Reinforcement Learning for Diffusion Policy Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:28:41.908138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:28:41.908138Z digest=sha256:a03cfaf3dd1b5b118f5170a76ef1ff033fb9c413379087ec3c7e08c6edbae2fc

Observation 26c18d38-9d6d-48d8-8cfc-eb8af3de670c · inbound

Habitizing Diffusion Planning for Efficient and Effective Decision Making cites this paper.

Habitizing Diffusion Planning for Efficient and Effective Decision Making Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:40:16.106279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:40:16.106279Z digest=sha256:6cf31b969b0be1c5d0396feb5c8eb3be96c06e255f2a39c4d9ce2047a5c4101a

Observation 42042865-f0ce-403b-a6f5-260f47240610 · inbound

Exploratory Diffusion Model for Unsupervised Reinforcement Learning cites this paper.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.374992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.374992Z digest=sha256:5ea59794076804f1a5e781a31588b048ced683267888ee55aaa118bab2b1d920

Observation cb614448-9894-4ff5-914d-2d667dd83192 · inbound

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion cites this paper.

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:09:51.254192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:09:51.254192Z digest=sha256:15df16f760381ff30ab5e998c3ac5ffed2dc5df73383f73798bcad19ccfcbd3e

Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · inbound

Flow-Based Policy for Online Reinforcement Learning cites this paper.

Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.346736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.346736Z digest=sha256:8cd528e1af3f9be30a0dfb6c2a488a6a4ad1f2c29097977fb5924247727b1fba

Observation 856d2f9c-7937-4712-9664-e6fde3acddc0 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.490841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:6df260b8319b3d5eb734d168b3b9b1074a475d3b7bb6bd4fff8dc2935c3459af

Observation f35fc052-b97b-48f5-a52d-0b7a802fb943 · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.226296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:0bacd5789a7b6097a3fa9a219aeb5976bee25b5b6afd233d7f1f73fa200491e3

Observation 307a410a-5ad7-4481-83ff-e91af1ade30c · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.645809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.645809Z digest=sha256:bddb7699297b0da276b9accee3e6d539d2ec04190b44d7b8e4283405b33c29bc

Observation a7318975-2a5e-4318-ba51-97ae2e6af7ab · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:09.827237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:09.827237Z digest=sha256:adaf652dcd4c10f1b587faadf6231cb5d356e9df58c9d3c2afa3f55ab60e352e

Observation 47c4faa7-5cc4-4405-9729-7c73890d20fd · inbound

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? cites this paper.

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:57:33.094604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T07:55:31.706717Z digest=sha256:1de156ab7220bc5fbf4d190405c8353e72ffba8f80f3fae21fbdff4b2ff53eed

Observation 65ae67e8-1760-4953-a541-4addab9581ab · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:6072a145d37f8048d4add63eba85b5da2989944343e431c70e54f72196d6db7d

Observation 7ba5bfc9-19c5-4378-97af-882b25215429 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.680725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:00132df99d74e319a60256f70bad20a1ac81a1b23a0f6d8dd3e4b78e184afa07

Observation cdef4ab6-d0d9-4680-9c2b-775c94be44c3 · inbound

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning cites this paper.

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:55.078354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T01:08:17.624506Z digest=sha256:1e1626b5816b89fac979bf9c136f7cf953a1796f06abc014d9db98374d44fa39

Observation 35a20015-daaf-4adc-843d-4d1405965043 · inbound

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing cites this paper.

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:18:59.880117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T20:15:44.030714Z digest=sha256:47f8f81fc66838f731baa7869aa19e2d2cc7c67fd5e204bee10669ee6fa58d6a

Observation 7782a234-9094-4d3b-9cda-6b1745b293cf · inbound

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization cites this paper.

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:01.348420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:46:27.179341Z digest=sha256:af9ddc3c0dcb96b4d7d94e69632ac6d0d5879f471e15a3d0b3c8553495f7721c

Observation da7e45d9-72c8-4795-bf24-8f1fcf9b86a8 · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.524120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:8d9fd88324f8fc2a1e258bdbfcf8ac2f78746de5cc32144c96c7beba3e46e0b9

Observation 1ba195ea-06a1-4644-a323-f277eed15319 · inbound

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance cites this paper.

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.577987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T07:21:56.382343Z digest=sha256:507802f0fd0665dec6f3c2f0ff41af5c6b69a23d840ad39cb5b3261833149047

Observation 1114145b-2ae2-47de-a7e7-4574ea8f6ad3 · inbound

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cites this paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:08.545297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:47:10.062975Z digest=sha256:c2060bc6a719d969f8a35d2cf6b6da7c5beeafe125c22d80b8673839b574f29f

Observation 6886e5a8-f91d-4a62-829d-c7acab63fbe4 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.939467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:214d4b1f697ea6a90c1ad404a6ff8fefa57f4c5b2ccd71cfb280430202871232

Observation 560ffcdb-62ca-4960-95ce-b0b3da32b78b · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:929c3d0570ebe84a1d3aaf362fcba7efe4a5e4be298471e6e46853e8543d986f

Observation 131fc597-58d5-4fb6-a993-89847e6b741a · inbound

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners cites this paper.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.881499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.881499Z digest=sha256:3981e594c04326e17fbaafd43a86eef2f7e318cf9ba2a35a4480a61f28f97f2f