Pith. sign in

Paper Citation Record · LEDGER

Dichotomous Diffusion Policy Optimization

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2601.00898.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.00898 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T13:18:33.153895Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:57:57.027963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.313058Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 170112b5-b494-410c-9f79-6a1c48ea97a6 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Dichotomous Diffusion Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.293564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.293564Z digest=sha256:0a531907b60f24f4dededb3f9452431899860b8b247d0afc6423c841418ee380

Observation 8c00f752-55d0-40ec-806e-661c329f41de · outbound

This paper cites an unresolved cited work.

Dichotomous Diffusion Policy Optimization Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.703700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.703700Z digest=sha256:aaacb1c1112b8a0088741bfcf4ae645dc34155b0ff3089db5ab83c3692bcb336

Observation 2031654f-4d5a-43fa-96c6-8b59febce21d · outbound

This paper cites URLB: unsupervised reinforcement learning benchmark.

Dichotomous Diffusion Policy Optimization URLB: unsupervised reinforcement learning benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.048860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.048860Z digest=sha256:3adb9c4a932bf291265f85e664d92e830d55b6933feaf8a0c80192ead9b4425f

Observation 908817a2-4a1b-4c0b-bb1d-1704207c72b4 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Dichotomous Diffusion Policy Optimization Aligning Text-to-Image Models using Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.168819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.168819Z digest=sha256:d5c26383586f541d98bb0603f4b4e913add1509d8ad55366a19042188ddab850

Observation f2aef2ba-53fd-42ec-947a-885899c4e013 · outbound

This paper cites Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation.

Dichotomous Diffusion Policy Optimization Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.372744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.372744Z digest=sha256:4c9edec716d9df159d85ed636bc8b0e2f36e362f064b3c26c3b091febf2bcc82

Observation 2cc52555-cdb8-42f2-b81b-bdbf991d49c7 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787,.

Dichotomous Diffusion Policy Optimization Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.517734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.517734Z digest=sha256:db24fb2e6a6793573103630fdeb5464b48fa45e862ffd3e9f5f20acb915b8752

Observation 80952e50-9b4a-4819-ae7c-7c9e2cb459d6 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Dichotomous Diffusion Policy Optimization Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.632270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.632270Z digest=sha256:d6ad035390d007abdd92ad1fc88a3eddb9c723127d91b8d37014d12e29a82d0f

Observation b87e6b12-99d4-48dc-a898-d84e4b66d094 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Dichotomous Diffusion Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.860519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.860519Z digest=sha256:6ca662d3692a713467545e25095908a6ec23536e9f6e343ed46b1f21eee31d35

Observation e9441e47-57b0-4326-8325-3213ee3b2f43 · outbound

This paper cites Offline RL With Realistic Datasets: Heteroskedasticity and Support Constraints.

Dichotomous Diffusion Policy Optimization Offline RL With Realistic Datasets: Heteroskedasticity and Support Constraints

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.372455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.372455Z digest=sha256:ca457242581ef49667cb9f23da809aa068e838b4522d504895a58582bd4dadea

Observation 31c56648-6bc6-4cd0-b248-0da035a596a4 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Dichotomous Diffusion Policy Optimization Score-based generative modeling through stochastic differential equations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.491372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.491372Z digest=sha256:2f7846362d73429f00df12c0bcf618cd3704720a769a24dfd432613f9b75f375

Observation 6606e6f5-aa24-4179-825c-6d5a9a526179 · outbound

This paper cites DeepMind Control Suite.

Dichotomous Diffusion Policy Optimization DeepMind Control Suite

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.635163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.635163Z digest=sha256:9e77d2f2573d3ce2a8b54b55019b5f34d2a875d64c866b9ee081ee71ba8d2c7c

Observation f66d2dca-4c80-467d-9ed9-11657ba88217 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Dichotomous Diffusion Policy Optimization Behavior Regularized Offline Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.749822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.749822Z digest=sha256:3a054c68293b07447d1f495484ed10d33079d3f3c37b41c57f28ece4c7bf2a32

Observation c6a1765c-5ccd-42f4-9528-f69f63c67c47 · outbound

This paper cites An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning.

Dichotomous Diffusion Policy Optimization An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.831475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.831475Z digest=sha256:82fb7bf6d43c71980b93b9a053771520d4812489ffeba86b7ecd276ae9dc40e7

Observation df66ff61-e2f3-4b79-8359-27727a485f0f · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Dichotomous Diffusion Policy Optimization Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.944989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.944989Z digest=sha256:7de55296e77c6f3233d1f571817d47ccc09f3037cd91d7c7fd38b76648d4ad2c

Observation f840cb6c-c722-457d-a8f1-40ddd6899cd2 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

Dichotomous Diffusion Policy Optimization X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.062486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.062486Z digest=sha256:62d0d84634bedb4f50a8d7b2f954143dfc32cb6b9a2c8f749473b6675ff85ecb

Observation 4b589c11-a8b7-4284-bb5b-69da1d431f68 · outbound

This paper cites net/forum?id=j5JvZCaDM0.

Dichotomous Diffusion Policy Optimization net/forum?id=j5JvZCaDM0

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.162485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.162485Z digest=sha256:2e8d9fbe603690587f879a5ea3ec9273fcf18fe636cc0c24f6cfb9ba29f4cd25

Observation d949d42e-0683-4cd2-b345-907bdb1cd473 · outbound

This paper cites We utilize datasets collected by unsuper- vised RL algorithmsRND(Burda et al.,.

Dichotomous Diffusion Policy Optimization We utilize datasets collected by unsuper- vised RL algorithmsRND(Burda et al.,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.456115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.456115Z digest=sha256:267876ec13ac48db01f6a7750ca7675d610a4bf43f4aec4fb6d354a5634e1f31

Observation b1225208-3e2a-4087-aeec-39b4545c57b7 · outbound

This paper cites For each environment, we use the full dataset with all transitions from each dataset.

Dichotomous Diffusion Policy Optimization For each environment, we use the full dataset with all transitions from each dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.587921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.587921Z digest=sha256:b0845927bd793de0aa857ea5584601a3f86325666864615ca5ce919d71102787

Observation 79130141-d2ae-411a-8d08-052f5edf3486 · outbound

This paper cites Comparing with other offline RL methods,DIPOLEachieves better performance after the full finetuning process.

Dichotomous Diffusion Policy Optimization Comparing with other offline RL methods,DIPOLEachieves better performance after the full finetuning process

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.863599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.863599Z digest=sha256:4ea478136c87e367c841c365cb2a0df967c0a5057b4a6b8bcfc4ba9715c6a4d4

Observation 2138d679-b785-46a5-b753-f573f20bde8a · outbound

This paper cites The visual input comprises images from Front, Front-Left, and Front-Right perspectives, while the language input consists of driving commands provided by the dataset.

Dichotomous Diffusion Policy Optimization The visual input comprises images from Front, Front-Left, and Front-Right perspectives, while the language input consists of driving commands provided by the dataset

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.982972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.982972Z digest=sha256:7fb53620c9485ae69d1207a13fce93a5cd382c81d98766e4751a9b801c43beeb

Observation 217aa1d5-bf67-4493-88bd-130c6f8b7428 · outbound

This paper cites F LIMITATION& DISCUSSION& FUTUREWORK Here, we discuss the limitations, potential solutions, and promising future directions of our work.

Dichotomous Diffusion Policy Optimization F LIMITATION& DISCUSSION& FUTUREWORK Here, we discuss the limitations, potential solutions, and promising future directions of our work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:33.153895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:33.153895Z digest=sha256:5f6315ed7073479b349fb5f1d04a86b3beb7c35ce258e3488d2b12b5a844bb08

Observation ea263ba4-10b8-4dcb-befa-5b9ba01f6c48 · outbound

This paper cites In this setting, a policy is a probability distribution of actions conditioned on a state.

Dichotomous Diffusion Policy Optimization In this setting, a policy is a probability distribution of actions conditioned on a state

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:32.303397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:32.303397Z digest=sha256:1370577b21de65df7029876eed0ed6e46964c888e1755fc7b6ce9bfc21606e32

Observation fcf1c14a-838f-44d9-8772-5a807c1a88ee · outbound

This paper cites Proximal Policy Optimization Algorithms.

Dichotomous Diffusion Policy Optimization Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.128457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.128457Z digest=sha256:7752ab1bfbbbf867a27cacf024852a0214fba44eecb7d3b0b728c9cb849275dd

Observation efb88e57-6e14-4553-a5b8-ecb693f63400 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Dichotomous Diffusion Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:31.278572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:31.278572Z digest=sha256:06b2c7a7e6bc14f7b51714c6585cbb57545de66c745ea06b291a93199fd74a3b

Observation 7f145fc9-693c-408e-a139-8cd5c5c9dbed · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Dichotomous Diffusion Policy Optimization IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.599405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.599405Z digest=sha256:b31191045d32325f8720274130ec0bc56ef73a42a4544ac496d6d4bc26464a54

Observation a52b2a05-751a-4ddf-9a94-ef46be016991 · outbound

This paper cites Aligning Text-to-Image Diffusion Models with Reward Backpropagation.

Dichotomous Diffusion Policy Optimization Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.948259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.948259Z digest=sha256:006681f780591db68f03199cb4d98d8796f67fceb43cb95c262de4e26a67d954

Observation 7dcf3c5b-1ab8-40e9-a991-1fc2ccf06cb0 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Dichotomous Diffusion Policy Optimization AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.763809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.763809Z digest=sha256:c9ac272481587eec4bcfc2c64b8ce1022db658031c0c956f96aa7a5a8286aa20

Observation cba2875a-4da8-40c8-be8f-2d6167778d44 · outbound

This paper cites Diffusion Guidance Is a Controllable Policy Improvement Operator.

Dichotomous Diffusion Policy Optimization Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.364170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.364170Z digest=sha256:9b3ed12053ebb9c7dea623160f3c70e954bf9e753ef6b9a9cad39533ec8ccf46

Observation ffb3f69d-20ef-4e86-aa12-e5221f77c9c2 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Dichotomous Diffusion Policy Optimization Classifier-Free Diffusion Guidance

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.776718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.776718Z digest=sha256:7bf22171062cc78a29635587ceb50724e0cafcd75e5dcea6c423552d1d07ff8e

Observation f2b836f5-ff26-4d78-98fc-22e0dca0a632 · outbound

This paper cites Discrete diffusion for reflective vision-language-action models in autonomous driving.arXiv preprint arXiv:2509.20109,.

Dichotomous Diffusion Policy Optimization Discrete diffusion for reflective vision-language-action models in autonomous driving.arXiv preprint arXiv:2509.20109,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:30.248074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:30.248074Z digest=sha256:910466937a281e01cf850d498b61eed06f22c569528653e87f4bd75008062c36

Observation 514c0580-bd6f-44c0-b9bc-9e550952c025 · outbound

This paper cites Rl with kl penalties is better viewed as bayesian inference.

Dichotomous Diffusion Policy Optimization Rl with kl penalties is better viewed as bayesian inference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.914723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.914723Z digest=sha256:5b7130942d6879ea532eaf5ee2e77bfdbeaa90ba0417c16b047c651e1a734840

Observation d48eeea0-34f4-450f-9e87-ec6985f949f4 · outbound

This paper cites Extreme q-learning: Maxent rl without entropy.

Dichotomous Diffusion Policy Optimization Extreme q-learning: Maxent rl without entropy

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.439176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.439176Z digest=sha256:e268119041e190928cda85646fa52962876a7a4ab956ee69270462918b93c34d

Pith citing papers

Observation 2494b19a-086f-42ac-9622-946ba18891ee · inbound

CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies cites this paper.

CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies Dichotomous Diffusion Policy Optimization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-20T03:19:14.643152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:20:55.364175Z digest=sha256:e18fee088a4c8c6b67ab69cfe081b299cf16e7cd9665240d9a0d5af90ca25c19

Observation 2103e74d-c4ca-4cb9-8434-139796fbb3f8 · inbound

World Value Models for Robotic Manipulation cites this paper.

World Value Models for Robotic Manipulation Dichotomous Diffusion Policy Optimization

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-20T03:19:14.643152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T23:33:50.854261Z digest=sha256:f450866d2d24b4eb5d5aaf5c9af4b9535c1b279edf8fcec8badc0b10e88ea1b1

Observation c44bc475-1f86-4cf5-814a-33156b0f5d9a · inbound

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL cites this paper.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dichotomous Diffusion Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.027963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.027963Z digest=sha256:3a8f9a721be900e7b98921d9d16186858ba82ff6d43607ba4f9078eea19eb9b5