Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T13:18:33.153895Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2601.00898.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T13:18:33.153895Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:57:57.027963Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:40:00.313058Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 170112b5-b494-410c-9f79-6a1c48ea97a6 · outbound
Dichotomous Diffusion Policy Optimization $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c00f752-55d0-40ec-806e-661c329f41de · outbound
Dichotomous Diffusion Policy Optimization Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2031654f-4d5a-43fa-96c6-8b59febce21d · outbound
Dichotomous Diffusion Policy Optimization URLB: unsupervised reinforcement learning benchmark
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908817a2-4a1b-4c0b-bb1d-1704207c72b4 · outbound
Dichotomous Diffusion Policy Optimization Aligning Text-to-Image Models using Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2aef2ba-53fd-42ec-947a-885899c4e013 · outbound
Dichotomous Diffusion Policy Optimization Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc52555-cdb8-42f2-b81b-bdbf991d49c7 · outbound
Dichotomous Diffusion Policy Optimization Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80952e50-9b4a-4819-ae7c-7c9e2cb459d6 · outbound
Dichotomous Diffusion Policy Optimization Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87e6b12-99d4-48dc-a898-d84e4b66d094 · outbound
Dichotomous Diffusion Policy Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9441e47-57b0-4326-8325-3213ee3b2f43 · outbound
Dichotomous Diffusion Policy Optimization Offline RL With Realistic Datasets: Heteroskedasticity and Support Constraints
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c56648-6bc6-4cd0-b248-0da035a596a4 · outbound
Dichotomous Diffusion Policy Optimization Score-based generative modeling through stochastic differential equations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6606e6f5-aa24-4179-825c-6d5a9a526179 · outbound
Dichotomous Diffusion Policy Optimization DeepMind Control Suite
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66d2dca-4c80-467d-9ed9-11657ba88217 · outbound
Dichotomous Diffusion Policy Optimization Behavior Regularized Offline Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a1765c-5ccd-42f4-9528-f69f63c67c47 · outbound
Dichotomous Diffusion Policy Optimization An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df66ff61-e2f3-4b79-8359-27727a485f0f · outbound
Dichotomous Diffusion Policy Optimization Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f840cb6c-c722-457d-a8f1-40ddd6899cd2 · outbound
Dichotomous Diffusion Policy Optimization X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b589c11-a8b7-4284-bb5b-69da1d431f68 · outbound
Dichotomous Diffusion Policy Optimization net/forum?id=j5JvZCaDM0
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d949d42e-0683-4cd2-b345-907bdb1cd473 · outbound
Dichotomous Diffusion Policy Optimization We utilize datasets collected by unsuper- vised RL algorithmsRND(Burda et al.,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1225208-3e2a-4087-aeec-39b4545c57b7 · outbound
Dichotomous Diffusion Policy Optimization For each environment, we use the full dataset with all transitions from each dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79130141-d2ae-411a-8d08-052f5edf3486 · outbound
Dichotomous Diffusion Policy Optimization Comparing with other offline RL methods,DIPOLEachieves better performance after the full finetuning process
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2138d679-b785-46a5-b753-f573f20bde8a · outbound
Dichotomous Diffusion Policy Optimization The visual input comprises images from Front, Front-Left, and Front-Right perspectives, while the language input consists of driving commands provided by the dataset
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217aa1d5-bf67-4493-88bd-130c6f8b7428 · outbound
Dichotomous Diffusion Policy Optimization F LIMITATION& DISCUSSION& FUTUREWORK Here, we discuss the limitations, potential solutions, and promising future directions of our work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea263ba4-10b8-4dcb-befa-5b9ba01f6c48 · outbound
Dichotomous Diffusion Policy Optimization In this setting, a policy is a probability distribution of actions conditioned on a state
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf1c14a-838f-44d9-8772-5a807c1a88ee · outbound
Dichotomous Diffusion Policy Optimization Proximal Policy Optimization Algorithms
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efb88e57-6e14-4553-a5b8-ecb693f63400 · outbound
Dichotomous Diffusion Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f145fc9-693c-408e-a139-8cd5c5c9dbed · outbound
Dichotomous Diffusion Policy Optimization IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a52b2a05-751a-4ddf-9a94-ef46be016991 · outbound
Dichotomous Diffusion Policy Optimization Aligning Text-to-Image Diffusion Models with Reward Backpropagation
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dcf3c5b-1ab8-40e9-a991-1fc2ccf06cb0 · outbound
Dichotomous Diffusion Policy Optimization AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba2875a-4da8-40c8-be8f-2d6167778d44 · outbound
Dichotomous Diffusion Policy Optimization Diffusion Guidance Is a Controllable Policy Improvement Operator
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb3f69d-20ef-4e86-aa12-e5221f77c9c2 · outbound
Dichotomous Diffusion Policy Optimization Classifier-Free Diffusion Guidance
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b836f5-ff26-4d78-98fc-22e0dca0a632 · outbound
Dichotomous Diffusion Policy Optimization Discrete diffusion for reflective vision-language-action models in autonomous driving.arXiv preprint arXiv:2509.20109,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 514c0580-bd6f-44c0-b9bc-9e550952c025 · outbound
Dichotomous Diffusion Policy Optimization Rl with kl penalties is better viewed as bayesian inference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48eeea0-34f4-450f-9e87-ec6985f949f4 · outbound
Dichotomous Diffusion Policy Optimization Extreme q-learning: Maxent rl without entropy
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2494b19a-086f-42ac-9622-946ba18891ee · inbound
CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies Dichotomous Diffusion Policy Optimization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2103e74d-c4ca-4cb9-8434-139796fbb3f8 · inbound
World Value Models for Robotic Manipulation Dichotomous Diffusion Policy Optimization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c44bc475-1f86-4cf5-814a-33156b0f5d9a · inbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dichotomous Diffusion Policy Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.