Pith. sign in

Paper Citation Record · LEDGER

DiPOD: Diffusion Policy Optimization without Drifting Apart

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2606.13795.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.13795 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:06:28.426709Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact22
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 960b7e7c-b54b-4f7f-bbee-41bc7b37dc36 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

DiPOD: Diffusion Policy Optimization without Drifting Apart GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.946121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:68bf842a674cb50f77856c53c03dd49f12d5ba0e844c21f4ce82948b14746a60

Observation a22fbede-cd96-4da3-a2c9-e8bb06fef7f0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

DiPOD: Diffusion Policy Optimization without Drifting Apart $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
malformed identifier
local_arxiv, observed 2026-07-03T14:18:22.928069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:6f574d0a12aa8069a9dfd8e24dc4541ec910e7e3987f3e5739da07cff39bd6c3

Observation dfaa3aec-2983-4e79-b3b5-12ef839f308b · outbound

This paper cites DFlash: Block Diffusion for Flash Speculative Decoding.

DiPOD: Diffusion Policy Optimization without Drifting Apart DFlash: Block Diffusion for Flash Speculative Decoding

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:18:22.982158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:52ef391fcab6409fceed053ae2d3e937bfed57ae88c6d237639bce9274dcfca0

Observation 107974ec-2f41-4bf2-a137-1ae970526c6f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DiPOD: Diffusion Policy Optimization without Drifting Apart Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.925727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:73ef6329b274fff40402bbe988dc8450ae9e8afa74ffe96a45be5f77f1dc45cd

Observation aa89b0eb-d069-4902-a6f1-8c0dbf805984 · outbound

This paper cites EXPO: Stable Reinforcement Learning with Expressive Policies.

DiPOD: Diffusion Policy Optimization without Drifting Apart EXPO: Stable Reinforcement Learning with Expressive Policies

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.979890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:98ef89a0b5f61569a59fbac80745e63f3b339e7254ac0c81daa95219f0901cd0

Observation f34234b6-1274-43aa-baf0-8d82665e6cc1 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

DiPOD: Diffusion Policy Optimization without Drifting Apart IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.965715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:0f9a3f04ed75eb4b851d3b2f10cd190cc78a619dea7f169530fc883d4e2b0423

Observation 7016376a-1395-4f43-a273-9b4b3513c35c · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart Imagen Video: High Definition Video Generation with Diffusion Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.968033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:4efd2433a3d143b7a567a89137c68f7605db30dfff2ed9ab3a17dc8c6928b292

Observation 2b29a6f0-c9f7-4fef-beb7-ffc3689111d0 · outbound

This paper cites Q-learning with Adjoint Matching.

DiPOD: Diffusion Policy Optimization without Drifting Apart Q-learning with Adjoint Matching

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.970338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:6029c72488e2e94e18a9617d8a12423966e97bc1e32f431c62b8808a73afa942

Observation c4e80b62-0b03-4115-ba8e-966dfc9001e4 · outbound

This paper cites Reinforcement Learning with Action Chunking.

DiPOD: Diffusion Policy Optimization without Drifting Apart Reinforcement Learning with Action Chunking

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.963367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:27aa3851b9f4f2a2130efb1849093c8bd2a8013dc2c763b71da99541e61d72cd

Observation bf5b6a4e-37b3-4eb4-9983-2349a56664a5 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

DiPOD: Diffusion Policy Optimization without Drifting Apart Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.948665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:007def3baca3fe551fcf5b64c18150a3377af60ebaac662ad9d1b557155ce53b

Observation 644ab85c-d687-4f57-9b97-ce1248b4bbbe · outbound

This paper cites Flow Matching Policy Gradients.

DiPOD: Diffusion Policy Optimization without Drifting Apart Flow Matching Policy Gradients

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.932973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:24bb4892ac27869e18e87cce5c773711aa02689fe62bac1c7c196ffec5559e62

Observation 2b9f8ef1-dd90-43e7-bad7-8bb15d3e9c48 · outbound

This paper cites Large Language Diffusion Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart Large Language Diffusion Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.961036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:f0bc2cade4a50d7fd36feaab87d86fdd780cf03851b7cd70af5aeaac9af00fe1

Observation 9f694d9c-246a-488f-b50c-04c35aba46ff · outbound

This paper cites Diffusion Policy Policy Optimization.

DiPOD: Diffusion Policy Optimization without Drifting Apart Diffusion Policy Policy Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.958694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:bddefd6e2501b8fd21af4a56c89094f8e5f75af63fcf17986db894cf6c89fc50

Observation 741582ff-527d-43e1-a9b0-30a634f53803 · outbound

This paper cites Proximal Policy Optimization Algorithms.

DiPOD: Diffusion Policy Optimization without Drifting Apart Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.977532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:e84bcf14c5462efd5989bba9784edf09bbacfb1a4f1e0d454dd9b30160bb75e2

Observation 7c3e0754-8852-4ef0-911e-0011faceacf9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.975130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:a59c33facdb78a8948d16d71f853034c3771fe1cf1adcc2a18aa4decd238ebaf

Observation 0490dc30-4fa1-46f9-9922-76ad4b0bb0d1 · outbound

This paper cites Parrot: Data-Driven Behavioral Priors for Reinforcement Learning.

DiPOD: Diffusion Policy Optimization without Drifting Apart Parrot: Data-Driven Behavioral Priors for Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.938679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:0b33d893ac02b024ce296a3db3f537dbe65d7d064c24ef113d6c3cc4ec3a14bd

Observation c428be7f-c122-4f48-b4a8-844c4293810c · outbound

This paper cites Denoising Diffusion Implicit Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart Denoising Diffusion Implicit Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.972624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:ef03421ec9b0c66fa0f968e8d3d4984eebcdb39f140d3b078fe7f5797d512aa3

Observation b47c0824-bb87-4ed6-b25d-44a113ac9537 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

DiPOD: Diffusion Policy Optimization without Drifting Apart Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.930438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:44948c18a2bc58e325933bd722532f80698b8b6591dd48317cf480e5be8bd7d8

Observation eb61d956-522b-4463-9064-0dda553ee36d · outbound

This paper cites wd1: Weighted policy optimization for reasoning in diffusion language models.arXiv preprint arXiv:2507.08838.

DiPOD: Diffusion Policy Optimization without Drifting Apart wd1: Weighted policy optimization for reasoning in diffusion language models.arXiv preprint arXiv:2507.08838

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.936071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:90818de8e958a5abaa7e36ad48727c9f4b6eee154e7e8cd9ba28e596edb2ee0a

Observation 7c67e3c6-2da7-41d7-bd19-863e76a4da89 · outbound

This paper cites SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.956407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:7804a3d35618958522f74ce3bc6ba8029659e2f300b46981b87e39ab79efd87c

Observation 705f7fe2-ec23-4009-b52d-3b27e673a867 · outbound

This paper cites Dream-Coder 7B: An Open Diffusion Language Model for Code.

DiPOD: Diffusion Policy Optimization without Drifting Apart Dream-Coder 7B: An Open Diffusion Language Model for Code

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.941284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:d6d484ffd9c12a3a46b7b0fa81532acf90e697e34d978136aeadfdd0cd22a9da

Observation 95866b58-e1d8-4d40-9303-159209c7b898 · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

DiPOD: Diffusion Policy Optimization without Drifting Apart MMaDA: Multimodal Large Diffusion Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.943777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:05a0638f649c6585a2a73dd7705c05a6fba85c77d9cf97e96017b0819e6f099c

Observation 32b5ea62-73c1-44c0-8797-b94df4ebc65b · outbound

This paper cites Flow policy gradients for robot control.

DiPOD: Diffusion Policy Optimization without Drifting Apart Flow policy gradients for robot control

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.951369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:e3542c5d787a0ddda3a0dfb48fab44450fa3036eb91c3559d58e16da50281b92

Observation 0172605a-8bac-4417-a5f9-dfbc82e844bb · outbound

This paper cites d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning.

DiPOD: Diffusion Policy Optimization without Drifting Apart d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.953989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:dcd60c088585ac496f26a5502e6efc4d8282593f64d6720ee88d1449273ade94

Observation f4a60010-8b8c-40e2-9770-f1d742b55638 · outbound

This paper cites an unresolved cited work.

DiPOD: Diffusion Policy Optimization without Drifting Apart Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:163dcd0ada1d94f1209c97218566a5655498e67f6eb90ccb589598341e74d95a

Observation ee9e2e89-74e5-42e6-b1c4-873b56624781 · outbound

This paper cites The environment transitions to state st+1 ∼P(·|s t, at), and gives the agent a reward rt =r(s t, at).

DiPOD: Diffusion Policy Optimization without Drifting Apart The environment transitions to state st+1 ∼P(·|s t, at), and gives the agent a reward rt =r(s t, at)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:bd75949be7a2a1ee17b2af5fa3b01ff6e4b21040874945d54e6780d0d86bf49d

Observation e22934e3-26cb-4881-a589-4491db5d6669 · outbound

This paper cites SPG [Wang et al., 2025] obtains a practical EUBO surrogate for masked dLLMs by exploiting the special absorbing-mask forward process.

DiPOD: Diffusion Policy Optimization without Drifting Apart SPG [Wang et al., 2025] obtains a practical EUBO surrogate for masked dLLMs by exploiting the special absorbing-mask forward process

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:9a4ef3a330a232a4c5901de36b57a7e00c32226b0a7ed1b1a2a476e20c2a5f78

Observation 74d6f7ac-cbfc-482d-a8f4-7f0b43bde1ca · outbound

This paper cites This proves Equation (26).

DiPOD: Diffusion Policy Optimization without Drifting Apart This proves Equation (26)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:12516aaf25e44afa67123b4153bd6dacf78439b15f49cbdf5a783cf4fa37fee0

Observation 340379b8-429d-46aa-937b-f9a255adaf79 · outbound

This paper cites Algorithm implementation.We implement FPO by updating the model parameters according to Equation (4), and SPG by updating the model parameters according to Equation (5).

DiPOD: Diffusion Policy Optimization without Drifting Apart Algorithm implementation.We implement FPO by updating the model parameters according to Equation (4), and SPG by updating the model parameters according to Equation (5)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:1c08699e1517407aafbeb8b89b62b415d6f75c1772a57ac2c0b16bf7131fa079

Observation a6a5c6ae-62b7-4327-a5df-1e088123639c · outbound

This paper cites Although SPG+DiPOD does not show a clear improvement at sequence lengths128 and 512, it is the only setting that reaches the near-100%regime in the zero-shot setting.

DiPOD: Diffusion Policy Optimization without Drifting Apart Although SPG+DiPOD does not show a clear improvement at sequence lengths128 and 512, it is the only setting that reaches the near-100%regime in the zero-shot setting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T07:06:28.426709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:ea2864549f306b8c5f41c7cf19b807208203a7e6c2f3e506466f2ad3389e18b8

Pith citing papers

No inbound Pith citation observations are available.