Pith. sign in

Paper Citation Record · LEDGER

Diffusion Policy Policy Optimization

As of 4 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 63 inbound Pith citation observations for arXiv:2409.00588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.00588 v3

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T08:48:14.776754Z

measured 163 of 163 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:03:03.126259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T09:59:44.479162Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact53
  • verified fuzzy25
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7fc3da64-cce0-4485-b8e8-fcd4bef622ea · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.108380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:13ec2ee1b2f18f922c0f9e520160100a13f4b35af45e6ae309f6ff8a9e73e730

Observation fa2e04cf-b36c-48a7-a144-549378347b5e · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.051574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:9e2ea0d9232bb74dad9c8ed895622bdb3bb6219bf30d8e7f2ecf30ccfb25448a

Observation e5c6c449-52ec-4354-bb77-d3ac64bdc76d · outbound

This paper cites Residual Reinforcement Learning from Demonstrations.

Diffusion Policy Policy Optimization Residual Reinforcement Learning from Demonstrations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.827071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:8b56586b6d98ca81cfc0b51edf5f11b8a3335d586ebdff720ea21c45a59553f1

Observation 13205ca0-4611-47f5-b046-f2e730159781 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.056278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:3787b2934d22e30d06cff34267332e505d729fbe23fa3bbbeb42a42787d8d989

Observation 037f6897-8e31-4219-b4ce-3dabb7d8ebd5 · outbound

This paper cites Ankile, A.

Diffusion Policy Policy Optimization Ankile, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.058689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:230e8ca1c1dc5531957daa63085a91a32b7231961f5dfa63481df2b18525474f

Observation 74c87f12-a7d2-44d5-a6ae-f284aaa2d218 · outbound

This paper cites From Imitation to Refinement -- Residual RL for Precise Assembly.

Diffusion Policy Policy Optimization From Imitation to Refinement -- Residual RL for Precise Assembly

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.831533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:d6d3e27d255bc0e86e301c5f22c9027b54cd894b8cf32d5c85606352a20aea9e

Observation 8e01698d-6a82-40f7-93f7-5329ffdc77dd · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.063573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:44e6635c5c2172fd0cc7284b61bb8bc478188e54eb9cccd519c696a003de88c7

Observation 0f0df64d-6db3-4e5e-b853-a4ace7053ec5 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.066058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:2dca6bdbf22d8bb28c42664a4f89f81b9e2d59f5f322f925b760a50811865ea2

Observation 89beb46f-a3fd-4a26-a0f2-2e66b5fd0bf5 · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Diffusion Policy Policy Optimization Training Diffusion Models with Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.835367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:9d3d734d87423458e75341b98521d6c8cb8a7c7fd99ba6756cfae8e980024d6b

Observation 4119c290-68af-456a-8dcf-fd4ca288b698 · outbound

This paper cites Block, A.

Diffusion Policy Policy Optimization Block, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.070168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:457b32c37e28936628c2fc1401ad073734b511ea125bf44111136939690d218c

Observation e65f2b57-bd45-499e-ac9d-0847cc47d7c1 · outbound

This paper cites OpenAI Gym.

Diffusion Policy Policy Optimization OpenAI Gym

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.838541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:c1b06f45ee6b44c8b28f21aa1c31ed857b86fee15969a8ee3a8c73dc5a0e210b

Observation b0bb4f68-9b28-4ddc-91c5-15129b1e15fe · outbound

This paper cites Brown, B.

Diffusion Policy Policy Optimization Brown, B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.074531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:576f8095244c879f5aa3f2534c3622cb3f1dd4ff14b17289373abe13041d4b8e

Observation 1c8aa3e9-9d7f-4469-9a53-d60a1890d5ba · outbound

This paper cites Bruce, M.

Diffusion Policy Policy Optimization Bruce, M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.076598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:0e8a85bb29506ed59ddf2749feada11a27b5ea987fba66af86e8f5ed67dfd36f

Observation 62b94b51-a62f-45a8-a2df-08784e9f42be · outbound

This paper cites Tutorial on Diffusion Models for Imaging and Vision.

Diffusion Policy Policy Optimization Tutorial on Diffusion Models for Imaging and Vision

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.841858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:124d3bc42142801fa5f2ef1be8e2a0fd411ee495a839db9812d8d0a7978f1cee

Observation 4f6303ee-95ea-4f23-b7ae-bcb49599f045 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Diffusion Policy Policy Optimization Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.845798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:8d6b96e76116dd0ece73a3e787302d299dda70e5a779248ea5a397c45f0aa979

Observation 9f6973f3-5d8a-48a4-bcfe-65f9796f460d · outbound

This paper cites Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling.

Diffusion Policy Policy Optimization Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.849919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:e6dc4294b79e901be787dfbc837395b8ebe91e94aab253f644a649a61016e016

Observation b2db49b4-bb19-4625-8133-0cdf3387f210 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.084397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:bd886507417d9584b5b6f16e36329997d69fbf6a35c7c95fd5ba3c158a75ccfc

Observation 2da8068c-b370-4e02-96eb-99988b3b96ab · outbound

This paper cites Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation.

Diffusion Policy Policy Optimization Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.853322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:c30ecdb4367bc9990fc3586342670dd723c4199731ed833a53d7fb3ce60e8663

Observation 3722d88f-12eb-4f8b-a70b-40ebb751f6ad · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.088267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:abad844f212d50d763470e4477907528185fe44f7473c0f73ab3ffa4ddad7d65

Observation ce1b9978-aeab-4a09-b4c7-f8e851a084f8 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.090407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:905181b459998c1ae4dbec7c1e67709c28bc4cc367c3d364af22850c6cf72129

Observation 9388e5ab-fe5b-4e74-859d-4f6db40909d0 · outbound

This paper cites Directly Fine-Tuning Diffusion Models on Differentiable Rewards.

Diffusion Policy Policy Optimization Directly Fine-Tuning Diffusion Models on Differentiable Rewards

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:11:32.190572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:d77f8a7c93c59a95f1edbf082d28d9df78b6aff11cd4eefefb12a02161b7ee69

Observation b5700202-7739-4162-9ac2-0a8b20910eda · outbound

This paper cites Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning.

Diffusion Policy Policy Optimization Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.860719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:bcd37c29b2364485b13b60c7c4c3255a7a070795ad40cfecde9516c985130fd2

Observation b6fd875f-1610-48a3-a25e-5fb03b57d5f1 · outbound

This paper cites Engstrom, A.

Diffusion Policy Policy Optimization Engstrom, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.096544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:e13498cfc038475e6dd1ca0d415df31a4e9fa16d59054e4c65acf74429c20ac7

Observation f64c0a8b-9824-4093-ade5-fc045b372b82 · outbound

This paper cites Optimizing DDPM Sampling with Shortcut Fine-Tuning.

Diffusion Policy Policy Optimization Optimizing DDPM Sampling with Shortcut Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.864965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:52245b8ee4fcd7172da3748158999379679ea6d316b0ade04c33ddad55ae6eea

Observation aca13e95-0d89-4088-b5cc-eb503d7cb804 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.100372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:f878daf18b5a9531026cdf678c03d496b686856e9fed06c0330d7aa7b1f5df00

Observation bfcfe2a8-0a6f-4dd2-90d5-a339d57c9847 · outbound

This paper cites Fefferman, S.

Diffusion Policy Policy Optimization Fefferman, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.102333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:43d86dee48465712df188c85fd16536507ac1d9f6e480c8fe3cde15eb2b8ac20

Observation cf9196a7-3b82-43da-b8bf-349fa9112a12 · outbound

This paper cites Florence, L.

Diffusion Policy Policy Optimization Florence, L

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.104451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:0e26cf3158bd133c0e164f894de5b3aa58f4e336655247a12d0dd38d471811b6

Observation 89c55f4c-1d9f-49da-a28c-a0353b44d3fb · outbound

This paper cites Florence, C.

Diffusion Policy Policy Optimization Florence, C

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.106536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:aa7d4fc73db8084a685944fc97f940729114f9a160e2447fde56e01d8ea0b7cd

Observation 5a6f5bcd-6534-4829-b82c-b488f681c20d · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Diffusion Policy Policy Optimization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.822907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:7a72ece8bc384c5df55586f44748b71914170190ab1d77798d2bffdb09c7ccfe

Observation c5e61fee-a515-4fbe-acc6-addb25fb0708 · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

Diffusion Policy Policy Optimization Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.868566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:02e7b4566ecf37f69d6e294715a7f0ddbc7513f60898cba661278ae6df678254

Observation 373d9164-a265-41d8-bc18-049d89df86fd · outbound

This paper cites Know Your Boundaries: The Necessity of Explicit Behavioral Cloning in Offline RL.

Diffusion Policy Policy Optimization Know Your Boundaries: The Necessity of Explicit Behavioral Cloning in Offline RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.872015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:62d3b8bb1b35dd1ec941395564d141fd403b5f86c873af335eed639d648a1fe6

Observation de2072f5-0ece-47fb-96a7-13ab1a87fafa · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

Diffusion Policy Policy Optimization Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.875918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:2095fe963bdc28f02281e80df8ee533141d4fb06db78dfd0dc1b20eb34dc7e67

Observation 7ff373f3-343a-4e3b-9590-1721b0eb3a3a · outbound

This paper cites Haarnoja, A.

Diffusion Policy Policy Optimization Haarnoja, A

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.116187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:425197e86e2ceadd2e709c05df1e7af2d9d2588d221891f29ad7e2893fa8ed1b

Observation 5e66e647-3eb5-497e-9e62-6408778c974a · outbound

This paper cites Teach a Robot to FISH: Versatile Imitation from One Minute of Demonstrations.

Diffusion Policy Policy Optimization Teach a Robot to FISH: Versatile Imitation from One Minute of Demonstrations

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.879677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:6e6be16598f251cd1f726167c19c71ec809dbdd7ab4a165a5aca821348f96e86

Observation 878d79bc-74fd-4246-b73e-ca6f977d5087 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Diffusion Policy Policy Optimization IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.883343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:4805015af67e38c168e64c1f7d74e32443dd0e379418d1b5e863c7f10af0c62f

Observation e5bbb4ca-f2b0-41e3-993f-5cf03bfac378 · outbound

This paper cites FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation.

Diffusion Policy Policy Optimization FurnitureBench: Reproducible Real-World Benchmark for Long-Horizon Complex Manipulation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.886574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:84b4ce12e27fb3657cfcbf4dc7cac751ca88083123ff8373fe80fa7a731ffe35

Observation d738a798-3bee-4f2a-b78e-a572ec5163a9 · outbound

This paper cites Hester, M.

Diffusion Policy Policy Optimization Hester, M

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.123589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:0473384b9c7995ae1fcb869eb7230af448060d3e7bd7e4a0b5480cf1995dcf6e

Observation 3ef94dfb-c94a-4e4b-831b-9aab0bc0d2f7 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.125418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:a9d1ea06f7d5f92d086b15ddf0b611f938e664cfc81d6ee6366180d108fef95f

Observation b1e2ee15-62d4-4ef8-b3cf-6fe58860d580 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Diffusion Policy Policy Optimization Imagen Video: High Definition Video Generation with Diffusion Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.889827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:44cab28e6f76e8a208304ed4d8fcfcd70f9c2845c4995d15ef5bdb43cc223be0

Observation e2f0e55a-ef98-4bc1-841c-b4476a8e8720 · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Diffusion Policy Policy Optimization Imitation Bootstrapped Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.893529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:6eb1b3aa8415a38c9ccd07d762568044415140634c9e8529e1f9364649eb2afa

Observation 3e2957bf-5e4d-4657-8beb-0dec2eddcacf · outbound

This paper cites Huang, R.

Diffusion Policy Policy Optimization Huang, R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.130870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:f868c10237e4c5f123ffc0fef2d45e75240e64285dc09513bdc4bf09ff961a30

Observation 1b4e7401-0600-4b2d-b788-612d3b0ecee6 · outbound

This paper cites Huang, L.

Diffusion Policy Policy Optimization Huang, L

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.132632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b9d586dc88983a65de4afafbd80043835ebeceb2d80102a5b5295840f10d86a2

Observation e4769549-a34a-4c94-8cf0-43f9a838bfda · outbound

This paper cites Hwangbo, J.

Diffusion Policy Policy Optimization Hwangbo, J

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.134367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ae0764f590784396115f662635d1594173d3942a4a78f12d334153007ad73b91

Observation 207af9cd-d9ab-4bb8-8f6f-30f1af284488 · outbound

This paper cites Policy-Guided Diffusion.

Diffusion Policy Policy Optimization Policy-Guided Diffusion

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.896638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:25ebbb63a80b721b5070d82f16b462c0ea791bf54e660496676138022225c919

Observation 19cbe77a-fd06-4feb-a8a4-7841946094ef · outbound

This paper cites Planning with Diffusion for Flexible Behavior Synthesis.

Diffusion Policy Policy Optimization Planning with Diffusion for Flexible Behavior Synthesis

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.899710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:cab6c4fe382d3593036f710f1c2241443d60218f9f6fa7f531d9e512508c0319

Observation 9c6c5695-e309-4556-879d-e509aa1df0cd · outbound

This paper cites Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human Demonstrations.

Diffusion Policy Policy Optimization Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human Demonstrations

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.903087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:738d91674c3b93095319f4159a1416af9a9d9d6f29408c137b25770bca18b2c7

Observation 897764b3-e58d-4ce1-82b0-fd7a1eca23ee · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.142239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:06a3e06d8fe824f4108065bf11f02f5cf122f02368291bb7de484c68b3bffc77

Observation 6fdd6061-686d-4869-a20f-04de5dd0d5d3 · outbound

This paper cites Kaufmann, L.

Diffusion Policy Policy Optimization Kaufmann, L

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.144052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b92554da9a205d93159c7f9252540f420961f590762c6358f552eb935a3fe4f5

Observation a3e21a21-e1bc-4461-82b7-ef25800d6f75 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

Diffusion Policy Policy Optimization DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.906115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:d2fb3d675950091cebb5d1e2a8a22e9b15f24427ba2e47335172f06e8ea929af

Observation 155b40ea-1331-4e01-91cb-ab31e395b473 · outbound

This paper cites Behavior Generation with Latent Actions.

Diffusion Policy Policy Optimization Behavior Generation with Latent Actions

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.931553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:eea5ee90c1d6b065b32bcf8cc3f86a91e437e0cf55c34bb5f65ade7201a139e0

Observation 1356edd8-6914-40d8-9d0b-b7e6692ce847 · outbound

This paper cites Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization.

Diffusion Policy Policy Optimization Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy Optimization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.935556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:bbefedcefd770174d69616261c22e532d44bb19ad2e9aa20a341a0e55d1be916

Observation 742849f3-951f-4432-a524-cfd741545fc4 · outbound

This paper cites Learning Active Task-Oriented Exploration Policies for Bridging the Sim-to-Real Gap.

Diffusion Policy Policy Optimization Learning Active Task-Oriented Exploration Policies for Bridging the Sim-to-Real Gap

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.938657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:71cb8bbaa76732aba604f0ffa4ded3ede90747833bc639714ec6de2c05071aff

Observation 1b302e06-57de-4a1b-afc3-b4a893556f3f · outbound

This paper cites AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners.

Diffusion Policy Policy Optimization AdaptDiffuser: Diffusion Models as Adaptive Self-evolving Planners

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.941618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:c6f981828382770e2385bbab8e6b59b6e6e983eaa0fab684739eaf115ccbac3c

Observation df6b903f-f4d1-481f-b74c-616082f4f03d · outbound

This paper cites Continuous control with deep reinforcement learning.

Diffusion Policy Policy Optimization Continuous control with deep reinforcement learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.944595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:9ff5c72cfc49d799a036ea3243c5a6a52b1ddfb268f7a167b40c32f24617b74e

Observation 0982a0eb-cce3-41b3-8c02-48d8a9a31d43 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.157444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:69491cf8d4876df6bed6756f11c8a68c4a09a4eb1a0efadb31534aee5b38ed09

Observation 604884e4-c53e-49c5-9bc8-199d48c873d9 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.159249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:8d91d3e0836b45615defba0f9f413bc175f66f9be762c3fb78e32807ff5cb6f7

Observation d991f1a3-6ec4-42a9-b978-4fe163b283cb · outbound

This paper cites SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning.

Diffusion Policy Policy Optimization SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.947648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:16d3410035ae17c5e31212f496fffed9efcd8632d5cb7d1a8f8d5eced9a8c8ca

Observation f0e52589-5d3b-4192-bf35-4835a00ef1d4 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.049191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b647a9e5d948aca92f16e4e457b5c1b67f89ef03aad74360e0aecfce8d6c9a5e

Observation f15a744a-7a90-4d55-bc57-ef1cc7e6f697 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Diffusion Policy Policy Optimization Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.951019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:455b2c39a09aa99d58f0d69e512e2c882905486d9f170fb683bd2f4afb59f305

Observation fa2403fc-9764-4ae2-89f9-24f751ce7ca1 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

Diffusion Policy Policy Optimization What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.954439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:e78fd205bd69de83846db88207ff95236c4f4d5fdbc0dbed6b9cfdcb9f65e14b

Observation 5e7fcb95-9787-46ce-9e07-40e410eec516 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Diffusion Policy Policy Optimization AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.957892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:105e5ff8e52f4c1d7ed739f0f05e1bd7358665f7b398507ef0cb00cbcdac5c42

Observation 36a48b8c-c2df-41d0-a54b-62274b621b84 · outbound

This paper cites Nakamoto, S.

Diffusion Policy Policy Optimization Nakamoto, S

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.072228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:fdc80a2ceb638288177ff251ed5a6dd8380f47b16cb79dac79ccf48096be35bf

Observation 01610e55-c2d7-408b-880e-25024554c854 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.078608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:00fc33c4c6625217a38b6f714b71c0fdd6010d4cd0c00a5f02c40c7a28c016ca

Observation b5f3fc1f-9d8f-40ce-a304-bfc341f8a582 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.080483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ccd78c9ae991bf96d345f859e16339645edecd7716bd8258eb8cfae4c8c85b24

Observation d16ff11d-0094-492a-b86c-7a6160eb0b50 · outbound

This paper cites Ouyang, J.

Diffusion Policy Policy Optimization Ouyang, J

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.082537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:33903fab031338711f607077d4d9c49c8bd7c0073458ba3094077c0d9487e186

Observation ba3952f7-0862-43f8-a735-ebde9901b913 · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

Diffusion Policy Policy Optimization Imitating Human Behaviour with Diffusion Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.960944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:b45a59bf166eafadb828531079bad45e02e10e55e07b379c5903ae782794c797

Observation 428dd748-3e80-4a20-8a92-146606bf9833 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Diffusion Policy Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.964693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:3bd081e9a3292decbb7acc96db2c8e0e4267445a67ac9bb49e264bd9add2c309

Observation 4a0b17cc-8989-40af-8bca-ee7a3b823343 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.094621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:76906ed359bd8a555369496d924e7f89776242703c475a4e2ad6d38227c7bced

Observation dc68bc04-70a0-4434-9196-8803b80c98ac · outbound

This paper cites Interpreting and Improving Diffusion Models from an Optimization Perspective.

Diffusion Policy Policy Optimization Interpreting and Improving Diffusion Models from an Optimization Perspective

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.967950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ccf4fb3676cb2fe58bbea832c9508e803d7e64004529fb22f88d4c12f93639a5

Observation b8f59e9e-3868-47c6-9d33-74f7baf4762d · outbound

This paper cites Peters and S.

Diffusion Policy Policy Optimization Peters and S

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.110529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:bf9d44a1eaa43dd13dbec57fed0f979f65aeeaf34a535a3415a7ebc768e04cf0

Observation 0551ad93-1d15-41f8-acd3-b3b5e711fac2 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.112446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:bf24f7dbb01f66944ff62cc0528cc637d0d766b9fbc2d9abb34f04ab4c681311

Observation cd5b7c29-02b9-421b-b43f-410a2df7d7db · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

Diffusion Policy Policy Optimization DreamFusion: Text-to-3D using 2D Diffusion

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.971526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:6f8131089c9f4ade21eb790dec0df2b219244bf20ad699255d303afaf554affd

Observation 5de135df-59cb-4786-8bbe-515a5f0d00fe · outbound

This paper cites Popova, O.

Diffusion Policy Policy Optimization Popova, O

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.118047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:2a91b27d1d703f6eae72acaa1b9c8e7f73d4968f0bd3f8a08cf4b086523fac49

Observation 16c179a2-3336-4052-94c8-4f91798adf1a · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Diffusion Policy Policy Optimization Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.975258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ed280ef73abfcfed8cbfe606cfffb79a6c1838ee67883167cc86d67d5f20c077

Observation 5e2d455d-2da9-46c3-9821-d7c39478237f · outbound

This paper cites Radford, J.

Diffusion Policy Policy Optimization Radford, J

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.121888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:1892773f698f08b57eb58fc355ccd6c2ebaf856c5bf96b242e2a9bd0884b4f8c

Observation 8bc30cf8-4c1a-4898-8881-acc51565a62e · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Diffusion Policy Policy Optimization Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.979202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:5657f13741dde7848342adebd7f21085f620234e42a8b9331c7b3886ee270dfc

Observation 499f6d8e-ba36-4e53-b0db-65f384514ed5 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.129159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:a3491ea2bbab27ec06a827ee5e272f024a90bf92d5466288cde2e6fa00bd17c8

Observation a3805bbf-4098-4eca-b228-8771291a13d8 · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

Diffusion Policy Policy Optimization Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.982840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:c42890fa974f60276b30eeec939792034882ba8a917bb6865dbb8e9610d07727

Observation 663e197e-a569-400a-8227-212c98c82aeb · outbound

This paper cites World Models via Policy-Guided Trajectory Diffusion.

Diffusion Policy Policy Optimization World Models via Policy-Guided Trajectory Diffusion

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.987049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:09db53746e7f3f58d7c42c1ce85d7bd07664f602c42a9a05f58ceab6d4b8cb9d

Observation 45bf2598-3577-4c36-a694-fb57716f0c1c · outbound

This paper cites Rombach, A.

Diffusion Policy Policy Optimization Rombach, A

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.140545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:d70227a7f4c83d48bb2451dc3726d90e3c8d8bfef07c6c740a43d99671382eaa

Observation 191b30f4-1b13-4b29-b0ad-f98b38730f58 · outbound

This paper cites Ronneberger, P.

Diffusion Policy Policy Optimization Ronneberger, P

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.145772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:35a8b06fdadb75c7392edf60c54413ba14e1c6f4842abfbee3e65e0123acbdef

Observation c3c90cbb-12ec-40fd-aade-ba21f527cd43 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.147724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ec4070bcdde67db8e4891648068cc48bf4b2229213c66d669af0bad0585b1b6e

Observation 7e7658b8-5ad6-4b95-98a5-b9ebda3c2b53 · outbound

This paper cites Simple and Effective Masked Diffusion Language Models.

Diffusion Policy Policy Optimization Simple and Effective Masked Diffusion Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:14.991270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:8f03395b1425a7ddf583cd7986df276c3ff9a7ea63d23683719b72e3b28b0d7c

Observation f69d719f-a5cc-474a-afbd-a2b2ae833406 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Diffusion Policy Policy Optimization High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.995362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:5cc42e2ef35117acea223723554e9835b2373ede96061a2a042fa8645fae297b

Observation b0ec06c1-2c5b-4209-879e-aa7929218867 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Diffusion Policy Policy Optimization Proximal Policy Optimization Algorithms

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:14.999602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:c2ba57a538402259237e497c05d02145be22e141123789eef007f758fb2995f8

Observation adeff572-726d-4666-86a6-d6c7e1315a17 · outbound

This paper cites Sohl-Dickstein, E.

Diffusion Policy Policy Optimization Sohl-Dickstein, E

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.155612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:1b013510d1fb985a9609c3b7067f95b9acdf61b9ee3e61af1d305a92794de6bb

Observation 5da1fee1-02d7-40db-bc17-97e7664aae02 · outbound

This paper cites Denoising Diffusion Implicit Models.

Diffusion Policy Policy Optimization Denoising Diffusion Implicit Models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:15.002755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:9e0152a5a7a8490b5987b0efcab240d52e0d7b1c6ac8826fa3a952b727c16e3a

Observation 4c0fb82d-3ff4-444d-9c07-425a9fa56ca4 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Diffusion Policy Policy Optimization Score-Based Generative Modeling through Stochastic Differential Equations

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:15.006507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:451591f1fc0a1cead2fae7c1cbebc1ceadee0c397d7d68bfe8a6305edcdee35b

Observation 5a73522d-5759-4371-b6b5-48190f546df5 · outbound

This paper cites NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration.

Diffusion Policy Policy Optimization NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.010016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:28186896114216c6bf855333768a0741f935082a27b375dbc9383826b707469d

Observation 356ceedf-4fe0-46fe-8e5f-362fcd95820a · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.068163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:9a8bf326672a40c4a78c9df229da1b11b508377e0c4ee852784ab45b267cdee2

Observation 713d4dd9-23cf-4335-80dd-543914033f93 · outbound

This paper cites an unresolved cited work.

Diffusion Policy Policy Optimization Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-16T08:48:15.086360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:35e3f6c87a2ee1db925054d6e367545b9fd0e915fe8f43a773a98e9b966165b3

Observation dfabe98e-6efe-48ad-8eef-0e80dba5264e · outbound

This paper cites Todorov, T.

Diffusion Policy Policy Optimization Todorov, T

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.092390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:390c32b0712afd7bf053e0be457120188f47f2f17e6d11dae8dba58e32f2d617

Observation ab7c209e-66ee-42e9-a4df-4d0b04c2c6fb · outbound

This paper cites Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation.

Diffusion Policy Policy Optimization Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.014213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:a82dc6f65da6192e0d8e766ae34e71da627bfd04a4de0f967a458e3d96143d0f

Observation 2e5a19b2-9bcb-4236-9478-35b9fe1b0f1a · outbound

This paper cites Vaswani, N.

Diffusion Policy Policy Optimization Vaswani, N

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.114286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:f0cb185761fc17fedb7647d42535e948557e415e6ebe6d4b64243affd7ade96a

Observation aba10201-d72d-4a14-8289-7536c0b95c42 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Diffusion Policy Policy Optimization Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:15.017380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:531a9bc21cdbe66d2e1001f6e384308b6fabab05455bec77105ab7d364fe7f01

Observation 90b479d3-48c4-489b-8680-f9d57dab6a0e · outbound

This paper cites Reasoning with Latent Diffusion in Offline Reinforcement Learning.

Diffusion Policy Policy Optimization Reasoning with Latent Diffusion in Offline Reinforcement Learning

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.020769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:d4bce882ee9bbfc0140b24b10c744e959e80c67ac65861d14ca9016a76da8423

Observation 47df5f88-d8fa-4abf-bd93-f6f80394c194 · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

Diffusion Policy Policy Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.024198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:e235a47ebf5700cce3e19bfeb97fb4365ca45a96b7ae9b66641a53d1b6749e24

Observation 315847f5-1fdf-4e79-b817-e2a1c2e27614 · outbound

This paper cites Wang and E.

Diffusion Policy Policy Optimization Wang and E

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T08:48:15.138538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:ffec8c90009a1b20e35e3b821248728e7ded631967a0bf0006f099b162f6e96f

Observation df7a9bab-7142-4db7-b37d-82d3dc65899e · outbound

This paper cites PoCo: Policy Composition from and for Heterogeneous Robot Learning.

Diffusion Policy Policy Optimization PoCo: Policy Composition from and for Heterogeneous Robot Learning

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.027836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:639b855cd018f92544d2a63fbeced3109a19c291d337ed51d3e25c64d662a8a0

Observation f6229189-5202-48aa-b716-dfd78823ce01 · outbound

This paper cites Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning.

Diffusion Policy Policy Optimization Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:48:15.032219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:48:14.776754Z digest=sha256:969fc1ca5eec6de7c2b1c578a327322651cf3cfa536cb2fdee1a7d01299ba0bf

Pith citing papers

Observation ba5ee6c1-f15f-4b12-aa3b-cca09a09d237 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization Diffusion Policy Policy Optimization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:02:41.883697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:e8c21ccd7711097b117593e76c723fa6aa142b75988b2a8a7c059ca717ae31aa

Observation 0c4db814-78d9-41be-9964-269ad20edbc9 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control Diffusion Policy Policy Optimization

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:e57311c7c8307c026872bd01957ddf004ba61762210e8402614bc8a2ec9a9336

Observation 47bc4938-94c5-4022-aabc-3252f79a485d · inbound

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving cites this paper.

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving Diffusion Policy Policy Optimization

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:36:24.319361Z digest=sha256:5e20e358819ef2fb2ec8ad8dc858769983e6da1c479d84cb36e4c34d71658c84

Observation 21565e65-1f4c-4528-880f-1c99c744359f · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Diffusion Policy Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:55:46.253261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:9096c57b416a0341e8e3d79098c418f7ebbe36f2903ad8a780de741340b42dd6

Observation 7f6d4f1a-cac3-4bb2-a553-46a2aed6c269 · inbound

Reinforcement Learning with Action Chunking cites this paper.

Reinforcement Learning with Action Chunking Diffusion Policy Policy Optimization

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:22:06.402197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T05:18:20.960945Z digest=sha256:e7f1d160e443c32b85fff062c92045a7d048a2b6988f6aa032f189b3c258174f

Observation a81dc09f-e3b4-45b2-b9b5-c3f9ace8201b · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Diffusion Policy Policy Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:12:05.270716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:69fa13ff9e2af239eac364a802841354cf2501a36c111acb44eddc110280a514

Observation 9ff6eb02-3981-4773-ab0e-31de15b9e121 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation Diffusion Policy Policy Optimization

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.521038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:a62bc52691b6a63749e03fc07150f64e4acc13a181d65c7e80f43ab68637994e

Observation 2826a469-44ff-4c82-a962-29060e3578a4 · inbound

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces cites this paper.

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces Diffusion Policy Policy Optimization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:15:38.728243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-21T21:14:54.053177Z digest=sha256:1d93b640b09a0ad9b7af4be902852cd9ed15e94c4364bf9b2e55c97ea050bfc9

Observation 578e1ed8-5be0-4d91-ae7e-10add8cfa854 · inbound

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning cites this paper.

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning Diffusion Policy Policy Optimization

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:58:55.143134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T02:55:11.396778Z digest=sha256:8ef876bb7af9c6fd68e0e4bc8b8e41871375bc9ba04ef93ff2647345f1c46ba7

Observation cc1aa7b9-c440-4da9-8d13-fac8edc2e47e · inbound

Training Diffusion Policies via Prior-Mapping Co-Evolution cites this paper.

Training Diffusion Policies via Prior-Mapping Co-Evolution Diffusion Policy Policy Optimization

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T19:03:03.126259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:03:03.126259Z digest=sha256:5bbaaf0d1be04ed0e4f2359df4ebf70fc90c4e6dc2414c0d8e4b87af4a0617cb

Observation 2500e2c8-52eb-424e-b955-b2c9b97629c8 · inbound

SuperFlow: Training Flow Matching Models with RL on the Fly cites this paper.

SuperFlow: Training Flow Matching Models with RL on the Fly Diffusion Policy Policy Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T15:58:43.680884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:58:43.680884Z digest=sha256:7a3f5f49a4cfad81f79573e29361573eb2e0f871c0cb7a59afbfce62f6aba673

Observation 11ceecf6-5757-4424-bb6f-8d9ac2889b82 · inbound

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation cites this paper.

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation Diffusion Policy Policy Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:23:16.097203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:23:16.097203Z digest=sha256:4eab5db6bfb24a3097c42b8ec6cc75f8525bb4d1ff559a8f104afbc7fa5cb4ac

Observation e9c845c4-cb15-40fc-ac0f-722adfefde5a · inbound

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? cites this paper.

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models? Diffusion Policy Policy Optimization

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:55:31.706717Z digest=sha256:85f5fd546d87dd5d6fb02beece7e54dfd70256c986e90c50bc03d8bc7a85f024

Observation da956c23-11d5-4ca9-9175-67ddb8e53664 · inbound

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training cites this paper.

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training Diffusion Policy Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:47:58.060371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:47:58.060371Z digest=sha256:6eb22d7a7c03c63b4dcf70c71b631804fa68786494a6687379595a2515ed57a6

Observation f02eaebd-af78-4b47-a2bd-066653aa463b · inbound

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection cites this paper.

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection Diffusion Policy Policy Optimization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:33:09.627731Z digest=sha256:ac8ae8235be3f593552eb578a9eddbf6e903f22f62a02443ccb231296e76a9cd

Observation 314312b9-a112-4e39-acc5-a801b507f8b3 · inbound

Space Syntax-guided Post-training for Residential Floor Plan Generation cites this paper.

Space Syntax-guided Post-training for Residential Floor Plan Generation Diffusion Policy Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T19:38:24.134463Z digest=sha256:c127d31b2297abaee1ed62d53dd00b89fa8517d8a82492f7f82a58bd0919ada1

Observation fd84c1b8-faff-46d5-8b92-c225d1c16a09 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? Diffusion Policy Policy Optimization

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:bfef670e27bd9e53ccda88822e083f4909310a1271905b298f0235592e7a9bfc

Observation e2a78e72-78c8-4cad-8e4e-acd0832556f6 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Diffusion Policy Policy Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:5978e08f4934a37e8cbf58c5568a404b637feacc4cc5529d4efe78cf07902886

Observation 5bde7aeb-4605-4e85-85ca-017a64f2e514 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Diffusion Policy Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:d2743ed5ce885d5959d568c74fa1c9455fde1586a34c0712c159e8f7d9d872f5

Observation 1849747a-3200-4c30-b7ed-693d3d74cb5e · inbound

You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector cites this paper.

You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector Diffusion Policy Policy Optimization

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T09:53:45.151347Z digest=sha256:907b26eba3b2b35971fd36e7b1c790d65b33bd8c580f838998a4b134b9cd1169

Observation 8a000a8c-5c2a-4e34-91c8-833fcf93f83e · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Diffusion Policy Policy Optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:5bc528e7809a62a0b22ec54a4693340a4faf7fc4268379596ca621b673e60699

Observation 8a6824f1-9b00-495e-99b3-66203b1639fc · inbound

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems cites this paper.

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems Diffusion Policy Policy Optimization

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T19:21:34.849729Z digest=sha256:3bba78ef6e5f8e6523c5d3f0c4518bdbf69f6bfa41827e8c16c90b55a701e87e

Observation 2d37fa2e-fa61-4baa-8f87-e7ae3f2b2e0a · inbound

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching cites this paper.

ScoRe-Flow: Complete Distributional Control via Score-Based Reinforcement Learning for Flow Matching Diffusion Policy Policy Optimization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:02:36.261596Z digest=sha256:101c2368e917573f9bd09ce43f80fcfc3a18ed5cefbce9455f55ba5a592786c8

Observation c7199917-fcdf-46e4-a5e2-1629a42e72cb · inbound

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement cites this paper.

StableIDM: Stabilizing Inverse Dynamics Model against Manipulator Truncation via Spatio-Temporal Refinement Diffusion Policy Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:41:45.903419Z digest=sha256:931aff8bd8417a2db6c0a805ed620c1417217c072ea01d9273d305677feec84f

Observation c8fbe686-b202-4159-8dc1-9a27a2657eda · inbound

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models cites this paper.

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models Diffusion Policy Policy Optimization

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T04:56:35.796962Z digest=sha256:88b8ba3c5f2dcecdcd86d4244561a367dec340c8930017abfa4de14eec1b5712

Observation 9ef6a3d8-ee54-40a8-9c76-ada06fbaa4cc · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Diffusion Policy Policy Optimization

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:de13911906b9d4dcdb237cc51e2001d0a285f1d0ffc91d761482dcc09ff3ab33

Observation 2af59ee1-3ad0-4010-935a-3ed5a7f4f754 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Diffusion Policy Policy Optimization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.649912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:6094a7757ab394a4ea7cbbd3780079dafe4954692c0d462d67f6976ea6cbf355

Observation 647b54c2-7905-4e05-a7ef-55c20243046c · inbound

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving cites this paper.

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving Diffusion Policy Policy Optimization

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T16:06:47.108382Z digest=sha256:a7fabb7768a4b1706fb68bccb9d600d6170fd02f4a9c910b0f92f001a03fda1b

Observation ec7fcb19-3c94-4d1e-a239-64f67249a527 · inbound

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving cites this paper.

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving Diffusion Policy Policy Optimization

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T01:48:36.105389Z digest=sha256:d541736db2412183f1ecd289f876f04e036dd77170ec97e636dc13b5ad0aa8dc

Observation 24450e5d-774f-45c1-a1a4-5e2f969605ef · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Diffusion Policy Policy Optimization

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T11:20:19.749415Z digest=sha256:8c4dc62b84283150852924042b69a5e98496dce53604b1694ff4a9f2033762b5

Observation ef41667e-d9f4-4a3f-aa51-e30b0168500d · inbound

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities cites this paper.

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities Diffusion Policy Policy Optimization

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T05:05:29.984355Z digest=sha256:9d8764e7bc71e5c5ac13865fe0facb96d753f58041f5b4f0af5b4adf05436b95

Observation b63fad67-3602-400c-bb82-0aba5003c69d · inbound

BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly cites this paper.

BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly Diffusion Policy Policy Optimization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:03:59.973863Z digest=sha256:8f75a97118582fcf77c8ee93316aabee53d1e521d288e92aad2812a4ac785e4c

Observation 14a2bfd0-4c76-4480-abf7-888cda1d2df1 · inbound

Driving Intents Amplify Planning-Oriented Reinforcement Learning cites this paper.

Driving Intents Amplify Planning-Oriented Reinforcement Learning Diffusion Policy Policy Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:50:14.602822Z digest=sha256:c05e9a063bb970ee961bc8c84d0090ec01211d41e97ad08e3765d984a38110eb

Observation 9f8ad26a-967f-4dd1-b434-c5f8df1f7fd6 · inbound

Driving Intents Amplify Planning-Oriented Reinforcement Learning cites this paper.

Driving Intents Amplify Planning-Oriented Reinforcement Learning Diffusion Policy Policy Optimization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T04:58:07.943615Z digest=sha256:d7ffe85e1ceb8a5586fcd6bb650f920b01bc1078754e8cf9691ac81a0f06a078

Observation a866d89e-0557-4c40-bdbd-0ee324810524 · inbound

EponaV2: Driving World Model with Comprehensive Future Reasoning cites this paper.

EponaV2: Driving World Model with Comprehensive Future Reasoning Diffusion Policy Policy Optimization

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:48:15.161771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:14:28.714494Z digest=sha256:b4258ad684f0c4723f93854669cf3ddb80514a8b9dbb2047b954a8d243cb4a1b

Observation 9bd20706-5a50-47fd-bd4d-6997829c682f · inbound

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing cites this paper.

Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing Diffusion Policy Policy Optimization

Reference 180

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:18:59.896329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T20:15:44.030714Z digest=sha256:3c4802dd51578d88072a84effa727a9239873e537f31e093f0b31cd71d1635d2

Observation fa0958ed-9273-4c4f-9083-cdd715b9a03c · inbound

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation cites this paper.

Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation Diffusion Policy Policy Optimization

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:08:21.214238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:03:45.869373Z digest=sha256:b46997fa03c92b360b59d12783e54dedfce8cb22a2a4468ebd15b243f19bb4e4

Observation 9fe43d4b-6039-4d0b-9355-5fb987ec830c · inbound

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control cites this paper.

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control Diffusion Policy Policy Optimization

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T09:44:05.614347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T09:43:07.291649Z digest=sha256:0a4c421c875b3db88f7ebe7642b6aab5155477c2705861cbdf37cc3455f3c548

Observation 5b537296-f68f-4ea4-8ca3-56e1747d33fc · inbound

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control cites this paper.

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control Diffusion Policy Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T16:17:15.396053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:17:15.396053Z digest=sha256:5556bbeca76be4a80f1cb526f6f0172bb79b9097053e128421e6148f81631dad

Observation 5a2c95a8-dfaa-4590-bfcb-9822ae1aaf70 · inbound

SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control cites this paper.

SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control Diffusion Policy Policy Optimization

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T02:40:14.646645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-25T02:38:14.542361Z digest=sha256:759aa6fc815189b4d455e8ea920e2c20caa9ec440a8fa863c38fbf3ef0d5ceeb

Observation dfd762db-b4d8-4092-912e-22a5e0bae85f · inbound

Score-Based One-step MeanFlow Policy Optimization cites this paper.

Score-Based One-step MeanFlow Policy Optimization Diffusion Policy Policy Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:26:39.040336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T05:21:52.748214Z digest=sha256:8df1f70ed92d13c662299985322a4804a042682ee914a1383986fe85a2ce5f1b

Observation 6de1a311-b8a7-4a12-b72a-95fde042912b · inbound

Adversarial Dual On-Policy Distillation from Expressive Teacher cites this paper.

Adversarial Dual On-Policy Distillation from Expressive Teacher Diffusion Policy Policy Optimization

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T19:03:51.530527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T18:55:23.777484Z digest=sha256:bc5e278c984e0c9b2d0b1e44d9c8e5e389808cd1e6dfdf5c8b3c283682ec87e0

Observation 4d781a1a-0f11-459a-bd6b-f7cc45f515cd · inbound

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models cites this paper.

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models Diffusion Policy Policy Optimization

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:16.332663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:44:53.969301Z digest=sha256:6a7b057482be24f27ef1d2af848510227677f4f3cbb95d7c2b18f64c0a128071

Observation 830c31ec-6ea0-45ac-8b4d-446e12a6f139 · inbound

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance cites this paper.

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Diffusion Policy Policy Optimization

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:23:12.560650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T07:21:56.382343Z digest=sha256:3121b34925c8267d3229798f15a9814df4ce649df8398152e9256e86b582c9cc

Observation af852c71-a5dc-4c63-ac5f-924ee87058f9 · inbound

Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies cites this paper.

Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies Diffusion Policy Policy Optimization

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:22:24.651659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:19:36.722892Z digest=sha256:e2c69d86a7b515c9cf139e07840773d061ec61954dd0744b61aa9067f051f90c

Observation 99f79ee2-17f2-4be2-8652-c6e9cc4b1b3b · inbound

L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation cites this paper.

L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation Diffusion Policy Policy Optimization

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:06:59.104806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T01:35:03.841581Z digest=sha256:b24518daf5a9154f3e800700da604e78b5ad4a18be71c8955b9953b25f169cae

Observation 1b8fe517-97fc-4df9-83de-54393ea24449 · inbound

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios cites this paper.

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios Diffusion Policy Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:27:08.810763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T22:47:10.062975Z digest=sha256:0ad865743d731603a5c09a2fdf9e1ed60cb1b4be4d8e9a253cb337808e014b5c

Observation 1b3afeba-f5de-4e9c-ba21-a2b48e66f082 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Diffusion Policy Policy Optimization

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:17:36.913056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:2c7b6a9a1d0896995a2920434128ba4789ebf5a18f3aca23a6a22dba5320160e

Observation 9f694d9c-246a-488f-b50c-04c35aba46ff · inbound

DiPOD: Diffusion Policy Optimization without Drifting Apart cites this paper.

DiPOD: Diffusion Policy Optimization without Drifting Apart Diffusion Policy Policy Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.958694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T07:06:28.426709Z digest=sha256:b7c16970c50afcdbff4ac382a58f498ccc38928911c471ffb40d8927c4039a74

Observation 89153b87-9389-418d-a987-6fc42e9c112d · inbound

Training and Evaluating Diffusion Policies with Long Context Lengths cites this paper.

Training and Evaluating Diffusion Policies with Long Context Lengths Diffusion Policy Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T13:49:10.591494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:49:10.591494Z digest=sha256:29ed0b09ebf3e7351987535ba908929530a1160773b6a1636ff0858781e55d39

Observation dd56ff9e-6c06-4a90-9f40-1ff194b3a13d · inbound

DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning cites this paper.

DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning Diffusion Policy Policy Optimization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:29:22.874635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T20:19:15.766523Z digest=sha256:a2a0de2fdc0d059eab3e952529e4565ab314680bc4afe0f397aa38d5a4959f76

Observation ff475354-aa11-47bc-b37e-f498c9ceca88 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Diffusion Policy Policy Optimization

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:59:44.497332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:504a3d5f8e8f56ecece52f784a3e0ca023ce6d50a0528858d3bdfa1e5d7bb06d

Observation 8dcde12d-1cad-4f93-b565-a567e860bd7c · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Diffusion Policy Policy Optimization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T18:25:58.752687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:baa61ec3af34989ad44f54aa3c7cf3ddca59638fd93a5fe8a34ab21519d35269

Observation 5bd0025f-7a96-4793-96fb-5b517d53088e · inbound

RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation cites this paper.

RoamFlow: Reinforcement-Aligned One-Step Action MeanFlow Policy for Image-Goal Navigation Diffusion Policy Policy Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.742141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T05:46:41.398250Z digest=sha256:e9824d97660bc4c7273e9dc22e4c95136f47facdc68e33e44a01695b3db47fc0

Observation f29df593-48d9-405d-b8ed-45309a366282 · inbound

FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement cites this paper.

FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement Diffusion Policy Policy Optimization

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:06:52.397510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T11:05:51.857152Z digest=sha256:1417081feb761851ecba3703fa66e4a400e59000ec092e1a0ba3baa908d010d8

Observation 27ba7173-50ca-4891-b91a-935619f026d8 · inbound

WorldSample: Closed-loop Real-robot RL with World Modelling cites this paper.

WorldSample: Closed-loop Real-robot RL with World Modelling Diffusion Policy Policy Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:58:02.429177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T10:57:40.128651Z digest=sha256:de3bcd3521b26677502f1d5a61562af867be9351f1a95238b721398b3802702b

Observation 0a61ad18-c644-4b46-b3eb-f0a62b6fc8de · inbound

PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers cites this paper.

PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers Diffusion Policy Policy Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T01:56:20.542865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:56:20.542865Z digest=sha256:ba0f441f3daaf0df4f36a59916a288f96fa5ef1fa10261987ff82a7ee0302cd6

Observation e527f006-6136-4ae1-b4ce-ed5878d95496 · inbound

Source-Lifted Flow Matching for Intervenable Multimodal Imitation cites this paper.

Source-Lifted Flow Matching for Intervenable Multimodal Imitation Diffusion Policy Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T13:30:25.329162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T13:30:25.329162Z digest=sha256:1b3bc5b6e1268cf58a9f72367b0c956359929ce3d13672e2a33fcc369eb73faf

Observation c643290a-e972-4984-97c1-b0b32927e796 · inbound

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer cites this paper.

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer Diffusion Policy Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T08:30:10.584105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T08:30:10.584105Z digest=sha256:fdbc018c5e6d03f44121bcce9286382c4ef49e662c1b2662cfef781373183524

Observation a346701a-1427-4467-94a4-0992f343157d · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Diffusion Policy Policy Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:44:57.681794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:44:57.681794Z digest=sha256:7e961de169fbf961232472c8fb55282996371a58749909b9b2ec08d785a0ccc6

Observation 61f6c394-e367-4202-bd4d-14ec0c057c0a · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Diffusion Policy Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:43.203548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:43.203548Z digest=sha256:7b54c5f2837b4aeec7be8598c415aea4139239ec078f89950198a45d8d9c9b6d

Observation a66dba47-35d2-44f2-9e09-d6587bece679 · inbound

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? cites this paper.

Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Diffusion Policy Policy Optimization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-30T11:06:23.633833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T11:06:23.633833Z digest=sha256:4b460da2c4cb11f4cad3543f275bda4c7173cb02ce92d76d4714093d5c28eeb9

Observation 9f13788d-107f-4628-ba61-208e54ae897a · inbound

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning cites this paper.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Diffusion Policy Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T03:34:04.794445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:34:04.794445Z digest=sha256:51ce2f50785a7325546a155d22eda32b81ed96c86975a3a86216f5151dd14ae6