Pith. sign in

Paper Citation Record · LEDGER

Learning a Diffusion Model Policy from Rewards via Q-Score Matching

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2312.11752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.11752 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:44.592129Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:17:36.937596Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 16c179a2-3336-4052-94c8-4f91798adf1a · inbound

Diffusion Policy Policy Optimization cites this paper.

Diffusion Policy Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.975258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:8decafa404501b8770a963490ba03c04a897585528aa91ec9316e62455e94f41

Observation 07e6a49b-df04-46df-a177-0ec39454bb38 · inbound

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation cites this paper.

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:48:30.730504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:48:30.730504Z digest=sha256:d0966f61d90da10c2d5241db6f65f547e98d03123805226083038b72318d637f

Observation b399cbd0-d316-495e-9e4b-ca909ce2e3a7 · inbound

Generative Diffusion Modeling: A Practical Handbook cites this paper.

Generative Diffusion Modeling: A Practical Handbook Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:49:35.453168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:49:35.453168Z digest=sha256:3ff9d75ffe872b31d0b1aa01691419b7c2d094288b3f8d823674c44c094fb9e1

Observation f5b3d552-8eee-4bcf-b728-92f8aaadc188 · inbound

Efficient Online Reinforcement Learning for Diffusion Policy cites this paper.

Efficient Online Reinforcement Learning for Diffusion Policy Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:28:41.908138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:28:41.908138Z digest=sha256:2afe365a052d2d8fcae9fa5158825cb96ca6e989f818329fa3e6bfe392d893c5

Observation 26c18d38-9d6d-48d8-8cfc-eb8af3de670c · inbound

Habitizing Diffusion Planning for Efficient and Effective Decision Making cites this paper.

Habitizing Diffusion Planning for Efficient and Effective Decision Making Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:40:16.106279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:40:16.106279Z digest=sha256:a5d23174da13772cd6651323bf1dcc89f2d7dbb0354d84d6a8904186a23da7ad

Observation 42042865-f0ce-403b-a6f5-260f47240610 · inbound

Exploratory Diffusion Model for Unsupervised Reinforcement Learning cites this paper.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.374992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.374992Z digest=sha256:982be10db22b3ad242a693ccb7eb36daff676dfe4dbccad619df1563089ee733

Observation cb614448-9894-4ff5-914d-2d667dd83192 · inbound

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion cites this paper.

Dual Control for Interactive Autonomous Merging with Model Predictive Diffusion Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:09:51.254192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:09:51.254192Z digest=sha256:ef7cf5eb5744f2ddee6c0986368d02382688cbe7f2838734d7546428cfeea1ae

Observation 2821fadf-5e2e-48bb-86da-83d8fda249a7 · inbound

Adaptive Diffusion Policy Optimization for Robotic Manipulation cites this paper.

Adaptive Diffusion Policy Optimization for Robotic Manipulation Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:44.592129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:00:44.592129Z digest=sha256:3a1cf313b99305142cad4ad698a2724925c104d95956a949a1c23d39d068723f

Observation 3724894e-8df1-46c4-b523-f13dc4af0a9a · inbound

Flow-Based Policy for Online Reinforcement Learning cites this paper.

Flow-Based Policy for Online Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:00.346736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:00.346736Z digest=sha256:058fa61f12625bb2b5e5336390fb84cf75281e927be2c4bc827bec5f3083eb01

Observation 856d2f9c-7937-4712-9664-e6fde3acddc0 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.490841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:794e012d348582376285fc633ce0f66f33ba667893dbb440c975c05338c846a2

Observation f35fc052-b97b-48f5-a52d-0b7a802fb943 · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.226296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:be7f382f932848b7da93a91bace9271c47f57d86ee3a7f3f60ac6cede4f015a3

Observation 307a410a-5ad7-4481-83ff-e91af1ade30c · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.645809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.645809Z digest=sha256:cc052926aefaaa30e260729e2cb9a2544ba0394533f5baa0df6e8abf74aa17b7

Observation a7318975-2a5e-4318-ba51-97ae2e6af7ab · inbound

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning cites this paper.

FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:09.827237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:44:09.827237Z digest=sha256:882daebc742845bc079d5c2b0ff2db29eaa5b9c22d230f2bf663facc968b3f17

Observation 47c4faa7-5cc4-4405-9729-7c73890d20fd · inbound

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? cites this paper.

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:57:33.094604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T07:55:31.706717Z digest=sha256:38d958fbbd68d8aacc2b9d8378b6e9f32830ab1070ca43fd8764d1f696012b9b

Observation 65ae67e8-1760-4953-a541-4addab9581ab · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:ca3c0d36ce9503631d2f72731830c3fc98dc152a6f38d97de75df09b1a516a7c

Observation 7ba5bfc9-19c5-4378-97af-882b25215429 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:20:25.680725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:e9f484315983798c30f9e83c38b90360a37ae5a1512c8f824752062a46fe1f60

Observation cdef4ab6-d0d9-4680-9c2b-775c94be44c3 · inbound

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning cites this paper.

Decentralized Diffusion Policy Learning for Enhanced Exploration in Cooperative Multi-agent Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:55.078354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:08:17.624506Z digest=sha256:67ae32e1436b38041ee237f96be263561c27cb1b1ed0961da8a88a383a442ae4

Observation 35a20015-daaf-4adc-843d-4d1405965043 · inbound

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing cites this paper.

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:18:59.880117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T20:15:44.030714Z digest=sha256:b823ea129ddf2c4c546c4e16c7d849c9a4d5f23a13915176e12bf16500c62cd1

Observation 7782a234-9094-4d3b-9cda-6b1745b293cf · inbound

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization cites this paper.

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:01.348420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:46:27.179341Z digest=sha256:a4248cc22ee89ebfc85952274db95ef697ea1451a5f9c1a184c47062a068d6ad

Observation da7e45d9-72c8-4795-bf24-8f1fcf9b86a8 · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.524120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:08950dbda177a794afe1a340e615f6594187b128a5033bcb5a92076adae951ae

Observation 1ba195ea-06a1-4644-a323-f277eed15319 · inbound

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance cites this paper.

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.577987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T07:21:56.382343Z digest=sha256:261ac2aece0ee02f4ceb4608a3f6d485713570425a90d8e0534f462962ef3e39

Observation 1114145b-2ae2-47de-a7e7-4574ea8f6ad3 · inbound

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cites this paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:08.545297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T22:47:10.062975Z digest=sha256:1204e2259c610e1b3e35236c739b68fd612bc768fc20677ef000ddeebba50fb1

Observation 6886e5a8-f91d-4a62-829d-c7acab63fbe4 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:17:36.939467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:1cfb984af16b7f2a4e7a488ae7c055af042600b863540f8e4803802b3843e79c

Observation 560ffcdb-62ca-4960-95ce-b0b3da32b78b · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:6665fb951c45ee6545a56b2b30ef496c324ad4da3da15d98fb2e9061c3bc550b

Observation 131fc597-58d5-4fb6-a993-89847e6b741a · inbound

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners cites this paper.

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:44:30.881499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:44:30.881499Z digest=sha256:f98cd1292af06bf9fa8a516d3da53d6d29288eb369d5040808ffebeb671aadc1