Pith. sign in

Paper Citation Record · LEDGER

Diffusion Guidance Is a Controllable Policy Improvement Operator

As of 22 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 35 inbound Pith citation observations for arXiv:2505.23458.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23458 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:30.265947Z

measured 114 of 114 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:38:31.344854Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.122193Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 274acc93-49ec-4872-82e3-7fc79ccf1802 · outbound

This paper cites Diffusion policies for out-of-distribution generalization in offline reinforcement learning.IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies for out-of-distribution generalization in offline reinforcement learning.IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.230721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:24.913040Z digest=sha256:8df905e6efaff8ff10efc9260d7511433fd447337eccd92fc6d07056d7faa3d3

Observation 0e25982c-fdbb-4d15-a09c-5f11796d28ca · outbound

This paper cites Building normalizing flows with stochastic interpolants.

Diffusion Guidance Is a Controllable Policy Improvement Operator Building normalizing flows with stochastic interpolants

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.216123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:24.978487Z digest=sha256:046c85b65a62857aae8cc09cc24803e7a9488afb88b29656aaf6efb8cbfe4f33

Observation 8c8a5edd-a8ff-4ab6-bfc9-6f2bd24dac69 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Diffusion Guidance Is a Controllable Policy Improvement Operator Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.201252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.066955Z digest=sha256:8823664e9bf7bc3fb1f7403167a14189d58c85185054ae1d41cfe956aa394fd6

Observation d631dc44-c4c3-4791-a139-0e2f163d7972 · outbound

This paper cites Hindsight experience replay.

Diffusion Guidance Is a Controllable Policy Improvement Operator Hindsight experience replay

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.186004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.131856Z digest=sha256:cbf126418d71ff57233946b612132f10ed1eafba88679093d511314d95a8eac7

Observation 9d261bb7-52e2-42b6-9e15-7cd904e73583 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Diffusion Guidance Is a Controllable Policy Improvement Operator $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.208106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.208106Z digest=sha256:dba8fdeca15f20d53524cddd718dc17439626c9ebef7ea4233e564b8f4d9d0c5

Observation 58bcf97e-f0ce-4eb1-80b2-b42e8b204c9b · outbound

This paper cites Zero-shot robotic manipulation with pretrained image-editing diffusion models.

Diffusion Guidance Is a Controllable Policy Improvement Operator Zero-shot robotic manipulation with pretrained image-editing diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.296803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.296803Z digest=sha256:26ecc2d11c9b02cf67eb7d20078c92017aaedd6f3a0dbd1754f00719b26fd32f

Observation 714c8e21-4f4a-42e1-88a9-710d62b6844c · outbound

This paper cites Whitney, Rajesh Ranganath, and Joan Bruna.

Diffusion Guidance Is a Controllable Policy Improvement Operator Whitney, Rajesh Ranganath, and Joan Bruna

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.162096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.405048Z digest=sha256:94e50c8bf27229b52ee8a25351bb15bb3bf3c3db2de4e9cdd7fa67bdf5442d2b

Observation c8d9265a-1c4c-48e5-9356-e219f334da01 · outbound

This paper cites Offline reinforcement learning via high-fidelity generative behavior modeling.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning via high-fidelity generative behavior modeling

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.145715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.509248Z digest=sha256:fb00434cf76f621dcad38af0d98414e6cdc4e91b477ed23ee6a16ae542e4df20

Observation 74bde7e8-c041-4700-92e2-603a020bc4e2 · outbound

This paper cites Score regularized policy optimization through diffusion behavior.

Diffusion Guidance Is a Controllable Policy Improvement Operator Score regularized policy optimization through diffusion behavior

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.131431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.589564Z digest=sha256:55e7c146f229a108c5f2946abc89b0930cb0f6057fb4f141a56af07bb1a6c432

Observation 802a466e-2092-4a69-b502-dbd8eef22ece · outbound

This paper cites Aligning diffusion behaviors with q-functions for efficient continuous control.

Diffusion Guidance Is a Controllable Policy Improvement Operator Aligning diffusion behaviors with q-functions for efficient continuous control

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.115183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.704932Z digest=sha256:c02f7111af9b831fcdbdb2884329f17734042ba5c5578bd79354fb790627d467

Observation 31dbe662-30bc-4e13-a288-4e83934ed286 · outbound

This paper cites Abbeel, A.

Diffusion Guidance Is a Controllable Policy Improvement Operator Abbeel, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.099654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.778783Z digest=sha256:ce8731b7b2675234e92d851ca47a5ff9684361f233b50fff959dbf86ecf3841d

Observation 2e5c6a4c-cd75-47ac-a79d-c4ac3e03c3fc · outbound

This paper cites Diffusion policies creating a trust region for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies creating a trust region for offline reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.084718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:25.891954Z digest=sha256:62c9c7afcfb5b1e31a9e0c9cbfb45ed138226fa5f26171fca0efbb588b2a6527

Observation f99072af-9eb1-449b-aa3c-70647f1bfc58 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy: Visuomotor policy learning via action diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.998739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.998739Z digest=sha256:254aef5ae88941d331a0f25e7797aec404a9d07140f7a010c88e7e51cd9c07de

Observation 642ed821-922b-4b5d-a86d-7d811cc1397f · outbound

This paper cites da Silva.

Diffusion Guidance Is a Controllable Policy Improvement Operator da Silva

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.059739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.110964Z digest=sha256:b38d08bf639d4e91ae5e657d801279bc1c6caf263b4577b68101fd6ee2daeb76

Observation 6660750a-a6d4-4d59-a499-5734479665b2 · outbound

This paper cites Using expectation-maximization for reinforcement learning.Neural Computation, 9:271–278, 1997.

Diffusion Guidance Is a Controllable Policy Improvement Operator Using expectation-maximization for reinforcement learning.Neural Computation, 9:271–278, 1997

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.044191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.195008Z digest=sha256:80b722048fd6d67bb37c8cf399ea324755d1fdedbab902434b887de495dd0619

Observation 6baffa69-bd7b-42ef-9032-355ef9e9716a · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-based reinforcement learning via q-weighted variational policy optimization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.028196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.273415Z digest=sha256:f781a3a722f233b3c547003db32680f342b0c78f3d59cca17d212f3834ce9a10

Observation d2ae4034-f701-45e1-b67f-ce913f910fe1 · outbound

This paper cites Consistency models as a rich and efficient policy class for reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Consistency models as a rich and efficient policy class for reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.012294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.379402Z digest=sha256:830402bbecb4d45c70f0e7a0c6dfded9bdd96ab72d6d66f461071147c3444fc4

Observation 3dfa2dba-d65f-4a2f-b469-1a53c1577de1 · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations (ICLR), 2022.

Diffusion Guidance Is a Controllable Policy Improvement Operator Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations (ICLR), 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.994233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.455926Z digest=sha256:3fbf47a4e7b4eecb6883c3862f49f853c77d4290ad91bc7d47c29a66cd60b59b

Observation c051a217-5a4e-4391-aacb-46e6b944f61f · outbound

This paper cites Imitating past successes can be very suboptimal.

Diffusion Guidance Is a Controllable Policy Improvement Operator Imitating past successes can be very suboptimal

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.977730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.558212Z digest=sha256:808ad1f3e3929d084b906e0bd7e1d7d043cf34b342650c5f98cf825dfb64b1bd

Observation 5256e09e-c9a8-4162-a405-424367c63127 · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive learning as goal-conditioned reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.961567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.631969Z digest=sha256:7b01b5bcd8ccf98fbaab8b1ef794a5ae3e30a6c52b072919c648cdf2e6078733

Observation ab7d341c-fbe0-4bb5-bcbe-8476cdff4e70 · outbound

This paper cites Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.945632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.695872Z digest=sha256:0c70c81315e14e23c08a22d7db0ddf118409454623bce47921c6c8c69518474c

Observation 69bb8e44-52f8-4eaf-8a65-28adba2ab38f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator A minimalist approach to offline reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.930897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.785311Z digest=sha256:5f6d29454ddd0864844131d4207a734cbf0bd812cb5219a37432c9b67329d083

Observation 8d237293-255e-4cc9-bb01-6abf487dfdd2 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Diffusion Guidance Is a Controllable Policy Improvement Operator Addressing function approximation error in actor-critic methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:26.868584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:26.868584Z digest=sha256:084a28b9c36410ea4291d65bfb8f1ec8ec24a7d92a10af26392c47e9407cb754

Observation a683e37b-e065-4cb8-a17c-da924fc5e915 · outbound

This paper cites Murphy, and Tim Salimans.

Diffusion Guidance Is a Controllable Policy Improvement Operator Murphy, and Tim Salimans

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.906617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:26.952030Z digest=sha256:f6f1f2a44b14b9e2f7a301f2fe45459acc90fc297748c2199aa5513e24b9981d

Observation 00d8e094-27f7-4188-ae32-abe1b12d04a9 · outbound

This paper cites Extreme q-learning: Maxent rl without entropy.

Diffusion Guidance Is a Controllable Policy Improvement Operator Extreme q-learning: Maxent rl without entropy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.891327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.012832Z digest=sha256:e8091beb35433d5b879dd5f56c1810d244103446fbce64971a0f7de52778a683

Observation 5a58fd3f-0340-408a-aa74-34bfd3a72932 · outbound

This paper cites Learning to reach goals via iterated supervised learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning to reach goals via iterated supervised learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.875520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.077763Z digest=sha256:a248f4a6ca56521a93dcf0548c14f58e3ffc8a904a28435cf5bf0130196c8fc3

Observation f936e271-bb74-451b-9a60-626cda99cec8 · outbound

This paper cites Closing the gap between td learning and supervised learning–a generalisation point of view.

Diffusion Guidance Is a Controllable Policy Improvement Operator Closing the gap between td learning and supervised learning–a generalisation point of view

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.860415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.159496Z digest=sha256:9046feda53e7b971b47f2f85c4a1f89b045d90f4a70153c0829eb4872cc2a2e1

Observation b890cd14-cc24-450f-8c7a-cb04512583e1 · outbound

This paper cites Explaining and harnessing adversarial examples.

Diffusion Guidance Is a Controllable Policy Improvement Operator Explaining and harnessing adversarial examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.219561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.219561Z digest=sha256:ee4ab37dca514658ca93d3822bf14800b65c50c1db06a8dec63addc1b4d73b63

Observation 0af6ac9e-9dba-4833-81fe-212a580d3a1a · outbound

This paper cites Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.835362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.306393Z digest=sha256:042b9f9a201d55323a6c50052cff5631712a03227815501e1d2e51a337676779

Observation a609deab-73a5-4a8b-8378-70e300f433dc · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Diffusion Guidance Is a Controllable Policy Improvement Operator IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.377097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.377097Z digest=sha256:82232cbeb72858e34a1af18f81c5abac5144bffdfa3ea910fea0748f537f3f42

Observation d80827a7-1c30-4b8b-b0e5-1a18224f6796 · outbound

This paper cites DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.429381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.429381Z digest=sha256:9cee5c49d0cedd7324ac47ee65eea2ea3f38ef70c10ebbdbf621f6609cb6748c

Observation 5b45d15d-891e-4072-840a-4d8ea3a004fd · outbound

This paper cites Aligniql: Policy alignment in implicit q-learning through constrained optimization.ArXiv, abs/2405.18187, 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Aligniql: Policy alignment in implicit q-learning through constrained optimization.ArXiv, abs/2405.18187, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.509049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.509049Z digest=sha256:4c56519e95ba44301aa0ea25a382d53716b1ef6cdf6d97d3c9522e673b8ad9d0

Observation 4cdc18df-65ac-4108-9745-750ac70b7148 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Diffusion Guidance Is a Controllable Policy Improvement Operator Gaussian Error Linear Units (GELUs)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.600669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.600669Z digest=sha256:f93c37b7f51534d41e05003ead92987d73cc0ee00acc0149438f357235105a72

Observation 53feb01d-6d0a-48da-8d91-cd9763944c6d · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diffusion Guidance Is a Controllable Policy Improvement Operator Classifier-Free Diffusion Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.693635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.693635Z digest=sha256:4ed1d7d3a84e425fe6596cdf30c04f51914bb80071d10f88d2c2db057dede916

Observation 2b83f3ae-92e2-49cb-86b3-f8b5cf6b2d18 · outbound

This paper cites Denoising diffusion probabilistic models.

Diffusion Guidance Is a Controllable Policy Improvement Operator Denoising diffusion probabilistic models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.818742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.756726Z digest=sha256:587f197abd49d950ab3731228476a4b7be8b8adc3044f42873a2d246f2a1b187

Observation 79375b08-a493-42d7-b055-4c6e4e25eb82 · outbound

This paper cites Tenenbaum, and Sergey Levine.

Diffusion Guidance Is a Controllable Policy Improvement Operator Tenenbaum, and Sergey Levine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.803927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.830267Z digest=sha256:dae0570c33376ec52438a0ffdcb88e5697bd70640d9bf1652a92149761a0d0ef

Observation 04dfc25a-5712-416c-ad4c-bb68502e31b6 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Efficient diffusion policies for offline reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.790011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:27.927436Z digest=sha256:82b7bb73d9c132437ca5c81815d0ed125a2e57d892037029d269f82670c7c196

Observation b0a23421-8d9a-4410-a78f-015dc9ea282a · outbound

This paper cites Kingma and Jimmy Ba.

Diffusion Guidance Is a Controllable Policy Improvement Operator Kingma and Jimmy Ba

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.003819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.003819Z digest=sha256:859a616d0a19aedc9f5e595b0c6259a87a9328767ec0f84d953f5c91363cd6e0

Observation 1f1d601c-2551-4ac9-8bc4-695c4a5526b3 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning with implicit q-learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.766144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.120329Z digest=sha256:57091312b777c32a3cc9faac31aabedb3b714c7710c4527116684b293935ee6f

Observation 19b2c0ff-aedb-4e00-8e0e-01f2ac6f42ab · outbound

This paper cites Advantage-conditioned diffusion: Offline rl via generalization.OpenReview, 2023.

Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-conditioned diffusion: Offline rl via generalization.OpenReview, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.751229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.181310Z digest=sha256:aa9b4c6faa55c0b4b8f95a96a5b9dcadb46d98fd73810b277b89d8926a996b14

Observation b68cdd13-872c-4b13-9aac-37ecd00af636 · outbound

This paper cites Reward-Conditioned Policies.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reward-Conditioned Policies

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.274947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.274947Z digest=sha256:993fb4455a98dc4479d890490974c65bc6e3f8f19ee82fa32ef11c0b12ab8eb5

Observation 8ab444d0-758f-4913-84b3-886758f587d9 · outbound

This paper cites Tucker, and Sergey Levine.

Diffusion Guidance Is a Controllable Policy Improvement Operator Tucker, and Sergey Levine

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.736331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.365220Z digest=sha256:ebd3232e1068ffb90eef0ae72c0976b74a31c09b4547e634affebc9446ecec38

Observation 54fb33ee-57a4-4ae5-ba20-ccd08657ac2d · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.720162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.446855Z digest=sha256:c84edef869e6e39350cba95676a312cef5183f2c00fbe96cddb9283b57631595

Observation f14ddcfc-602d-4fa8-96ae-14e7d2bddbfb · outbound

This paper cites Lillicrap, Jonathan J.

Diffusion Guidance Is a Controllable Policy Improvement Operator Lillicrap, Jonathan J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.705213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.512443Z digest=sha256:59e5e1e75dd465162d18d9784de518174e790b292f6d978c055f82e59bd87604

Observation c7778eff-81cd-448d-adc2-3761d12556a8 · outbound

This paper cites Flow matching for generative modeling.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow matching for generative modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.536332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.536332Z digest=sha256:cac1633582cd37216eee4939ec6f242bee8a7ee4b93709c2572829a57f371145

Observation 40be38c3-64eb-46d2-be15-b5110bc091d2 · outbound

This paper cites Flow Matching Guide and Code.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow Matching Guide and Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.549452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.549452Z digest=sha256:6ac1aadbbf394ca9c66eb43c0f1c92c81e33c1f7b698ebb84143682783457dc0

Observation c15fd0fc-e658-40a5-9f74-445ad8470372 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.554370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.554370Z digest=sha256:611e0e88e264f458822ca3315ce68a4721ffffdc0223fdea0bffafc28f06c955

Observation eb4ac65c-1a4a-45fd-89c7-228952e56585 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.672148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.574898Z digest=sha256:17156709b13ac3807a1bb9f473037c6c14db78370319d53c6dcdc63bbefbd554

Observation 186a8a5e-bffe-49d5-8301-714d5d616d02 · outbound

This paper cites Learning latent plans from play.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning latent plans from play

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.653436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.653436Z digest=sha256:21cddeab726a617913a149baeaf3318859d0c896f59e1fa84c98b1d92ba9078e

Observation 8b39392f-0dca-4393-9244-8e326a2cd847 · outbound

This paper cites Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.647886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.713236Z digest=sha256:d3413356fa668e09d9ec9f762067716340b6b896e9d93f544bfe330e9af518b5

Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.792421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.792421Z digest=sha256:e85d3c699820035a16285ddb34b84d7381c2763cf94df418d728306d78a5bee7

Observation ad88bfd8-f5c4-412d-b5d6-4f3426b19802 · outbound

This paper cites Mish: A self regularized non-monotonic activation function.

Diffusion Guidance Is a Controllable Policy Improvement Operator Mish: A self regularized non-monotonic activation function

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.633448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:28.875496Z digest=sha256:877284002b2494e697acdd96be8a9caa3576e00d27b13c5b3161702579bf429c

Observation dc9e6f52-cd4f-4fc8-b454-080ea9473496 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Diffusion Guidance Is a Controllable Policy Improvement Operator AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.945562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.945562Z digest=sha256:4e7e7856f2e0fc2a1ddf894950fd76c61091e7a1eb91713df3ceb5019dcee975

Observation 63304ea1-0b86-4062-ae8a-489c57bf65e0 · outbound

This paper cites Anti-exploration by random network distillation.

Diffusion Guidance Is a Controllable Policy Improvement Operator Anti-exploration by random network distillation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.618102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.026219Z digest=sha256:8ffc7e9e7ea93396525ca2a761646c2a658d474fe4d8a733aedec654eb8d669c

Observation aff38423-441a-48b5-a61c-d9a8530723e8 · outbound

This paper cites Is value learning really the main bottleneck in offline rl? InNeural Information Processing Systems (NeurIPS), 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Is value learning really the main bottleneck in offline rl? InNeural Information Processing Systems (NeurIPS), 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.602591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.036298Z digest=sha256:eb396e43797bd6e803de73b710c5068fcb7bf821fcac148ec4d358a45a53e181

Observation e8f0d878-ac34-4dda-bbd3-57460468c85f · outbound

This paper cites Ogbench: Benchmarking offline goal-conditioned rl.

Diffusion Guidance Is a Controllable Policy Improvement Operator Ogbench: Benchmarking offline goal-conditioned rl

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.040779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.040779Z digest=sha256:350e51843909a2870fd839014960ea2bbac48f91876acc7c5d195e06d12c88af

Observation 47c62e4a-7ec3-413d-b363-d7a933bb5ccd · outbound

This paper cites Flow q-learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow q-learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.577999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.065263Z digest=sha256:bcb39039e84b85994738b825f86c00eb0205fdad8abac94974c46c592d1d36ae

Observation 7c802dd2-e7f0-4e3a-a7c6-b8d2d3ce315d · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.148020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.148020Z digest=sha256:8f6e203ad0ab7c06cc66e3a73e13df65ec0b0bc49114d234a1846ebf368cf386

Observation 2a6fe649-1620-49cb-9503-9196a9bba391 · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reinforcement learning by reward-weighted regression for operational space control

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.562100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.220323Z digest=sha256:e5e5653fa413ad52e1630e3662fe55d81b6d671981c0ced26f99d66b10d75e3e

Observation ea3880fc-0749-47e5-bbe3-6c9b3a1ad795 · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning a diffusion model policy from rewards via q-score matching

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.547890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.278182Z digest=sha256:6454845ee6e05bef073539dc35f550e01d0cc763df70ac33fd2bad479f2d35ae

Observation 31f9eb45-0ce3-4031-be9e-a5a4fe774f09 · outbound

This paper cites Diffusion policy policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy policy optimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.534055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.356106Z digest=sha256:c5b4a7821248ba6c579380153416669d6a0ce146e072726141703c31ee47b616

Observation ddb9c88b-0d86-4d36-9e6d-b96d693575d6 · outbound

This paper cites Trust region policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Trust region policy optimization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.520084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.437576Z digest=sha256:b70ebe1cb331a1bfc229f35003ac10ef5e52ae03a9e31596704a1af5fd4f1e07

Observation 5794becc-5d9e-4908-a0c2-dc1fb1b1ee4d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Diffusion Guidance Is a Controllable Policy Improvement Operator Proximal Policy Optimization Algorithms

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.495573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.495573Z digest=sha256:116d1dddab6cf317c78605e4f8ffc14be9065a66805006071818de344456299c

Observation f63e00e8-fe76-4f37-977e-baef64f090e5 · outbound

This paper cites Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum.

Diffusion Guidance Is a Controllable Policy Improvement Operator Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.506492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.534334Z digest=sha256:4ec78541354caf397f0f8ec541d483413fdd1a664c209b0374c4125c96625fb4

Observation d8c70b07-4674-4b80-ab56-1f72b7ca04a8 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Diffusion Guidance Is a Controllable Policy Improvement Operator Deep unsupervised learning using nonequilibrium thermodynamics

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.538385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.538385Z digest=sha256:a42ac3c3d61d14e389df1c31be5ce1092b19bd0bbdbd98a0ff177fdc6b2bc46e

Observation ae689195-a547-4cfa-a0ef-fcc6fba41660 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Diffusion Guidance Is a Controllable Policy Improvement Operator Generative modeling by estimating gradients of the data distribution

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.483749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.544585Z digest=sha256:a768592d0837936cf2fc56c93f4c0e66930c34a2f9812c7865126e184d3d1e20

Observation 5871684f-043c-42a1-a6d1-2b1e394593d7 · outbound

This paper cites Sutton and Andrew G.

Diffusion Guidance Is a Controllable Policy Improvement Operator Sutton and Andrew G

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.469657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.612569Z digest=sha256:d096c100b5603f2bf32e6252eaae413d20beaa82ef45a015f67110da74727c6d

Observation 6b9578f0-d2a8-466d-8172-62933aaaca77 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy gradient methods for reinforcement learning with function approximation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.442908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.689334Z digest=sha256:3c743f15e3216ec55e0d02a7fbd8be9ad038141738599b8ac1100f679e586b9c

Observation e82727e7-5032-43ee-9501-5bc00cdd03b7 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Revisiting the minimalist approach to offline reinforcement learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.760477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.760477Z digest=sha256:e2dcb7d1e8df2f21788e5baead76b9913c74ef47b402631008aedef6a4715b9b

Observation 8204b323-f5b4-4fe3-a956-2fbf79d60459 · outbound

This paper cites Learning one representation to optimize all rewards.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning one representation to optimize all rewards

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.328623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.799586Z digest=sha256:8b8a814658d0b1c9793373ed3eabec57bc10d595a870d282a21ab0d27d70db68

Observation 44ea1e02-a513-4732-a729-ed5ce1f0510c · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.182390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.858233Z digest=sha256:74e2ee5937077e5693c4b7a471947ba66e93e8c07dad3c630d4d1346438a7a20

Observation 594f885e-0130-465b-a3d4-70f16a80e98d · outbound

This paper cites Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.106882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:29.933259Z digest=sha256:6289c0132658f3c14baf5e8d7533ed5fd52ad791a730de889aedd8a91847eb58

Observation e7adf455-ba68-41b8-ab50-8952b9fb234c · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Behavior Regularized Offline Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.993903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.993903Z digest=sha256:8ec980188da51c5ac18ac8f1eab9a5d4c41b98f1b7c21671dbae2212def67d24

Observation afcc1f24-a588-4efc-ba1a-105e0309a826 · outbound

This paper cites Offline rl with no ood actions: In-sample learning via implicit value regularization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline rl with no ood actions: In-sample learning via implicit value regularization

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.997347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:30.055766Z digest=sha256:f65ac43dc000106c067903aa1a663fa28ea606285296557dee80cbeecd2079f8

Observation 481580db-6b44-4321-8133-98cbe73ed38e · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Diffusion Guidance Is a Controllable Policy Improvement Operator Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.876244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:30.064184Z digest=sha256:f479907d17e0bf63b1f9ea43aeb347bed9f4c8c96d8f8093e6e4143540cc0173

Observation 749c8f5a-2bf8-4433-b336-ba5affe8f977 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:30.069061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:30.069061Z digest=sha256:6ad1df6327f35107aa47792cc051b942cf953e652f2b8bb34311b66442c6be99

Observation 062be4fa-d7d1-4753-afa9-7000e4ee3e37 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:30.096727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:30.096727Z digest=sha256:1fa9c4e86363eea9648cfc6b4aa6ab0b0028f82f0dd633ba61c154005fce5264

Observation b7214780-13b5-4915-a75a-6de01e8e131c · outbound

This paper cites Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.703344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:30.159233Z digest=sha256:35e2654a86e05dc0935608e992d58ef1bc413aecfa2bbd40f020b323de5fc87c

Observation fd15f4a3-d5a5-42db-ba9d-5a09e4d6c2c5 · outbound

This paper cites Energy-weighted flow matching for offline reinforce- ment learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Energy-weighted flow matching for offline reinforce- ment learning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.673893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:56:30.265947Z digest=sha256:548502fd9b24562a9a9a7733f6629b4225bd7343a846cd80699541e99ab778b4

Pith citing papers

Observation 56e107ec-6e8c-4c85-a1b2-0b52646da785 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.223628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.223628Z digest=sha256:91fc632b27482362ab0c759a74be430a4d975df0a7772531999b42c3487b5f38

Observation e09793a3-be45-46dd-ab6f-0509b118ac7c · inbound

DiffusionNFT: Online Diffusion Reinforcement with Forward Process cites this paper.

DiffusionNFT: Online Diffusion Reinforcement with Forward Process Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:54:30.998401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T16:54:30.953199Z digest=sha256:f3cd2ed611af025e0e8af00580af303423cddc87445591b2c12ec794417adf52

Observation c674f213-37de-4d0d-bacb-bec5bad07779 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.250729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:058749bd7e409e012ae6c9cb6e570aff693277854d4895be40bcff4893b3c4f8

Observation 21705f32-f9ce-49ff-a85b-0490e878683c · inbound

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator cites this paper.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.374560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.374560Z digest=sha256:a4f2b9e6e755852b8e10060eca742a8d7e6dd46799b4f89bfc1025062f716641

Observation cba2875a-4da8-40c8-be8f-2d6167778d44 · inbound

Dichotomous Diffusion Policy Optimization cites this paper.

Dichotomous Diffusion Policy Optimization Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.364170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.364170Z digest=sha256:b4f8936c6567e6618a61cff54ad5db635c040fd02305fe607c5a30d60eb3b5bd

Observation 1b108cbb-24e0-4cb4-a360-d36d9a30a11f · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:30:31.938553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:ce6c0c6d539fdf39aaf64aa59d5630911c1f5db6c9b9807b5854138c6cf77983

Observation c50286fe-3ad8-49ce-8090-0f3f49962bf5 · inbound

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training cites this paper.

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:47:54.752480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:47:54.752480Z digest=sha256:f6df9c20c9bf75868793c1661df80f7fe642224843fa5fc9539899de3bf52373

Observation 6feff226-b30e-4dc5-ad74-e425c919bf0f · inbound

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control cites this paper.

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:06:43.044814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:05:39.797848Z digest=sha256:3052e16dc82d2ee6077616bf2c6f151677182752cebead57a0af0a9a636ee197

Observation 9dae9f6f-7c72-481c-868b-5f8291ec48e8 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:47c7a2224e29fa269b0952bbe5e0690c2bf415eaac02017bf171ae7a7d4fdffb

Observation 54c53f7d-15fd-40a5-b96b-befcb10abac6 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:81b6eb345b5cb7117296796473e6dead2bd5411a25b7e8ea1182c230534a5fa6

Observation 99916ee1-1754-44d0-ac18-3fa2ad6c56d5 · inbound

Update-Free On-Policy Steering via Verifiers cites this paper.

Update-Free On-Policy Steering via Verifiers Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:316a03413ca1534e88df23661d4548ab4c45bcb3a9990f64163caf9344fda94b

Observation cc3c10a8-86bf-4b35-95a9-20bf1f2df124 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:25:59.160153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:e8650755a7e7866938b3abb25b8b64d06e311fa2f72eda7ac837aeab7d65a83b

Observation b7a92b01-d2b5-4d81-a199-b009fea7e531 · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:149eaf659dff712fa906eb42060c6c28f2023f71681c1a1dd41f9267a862700d

Observation 116dcf5c-f634-4dfc-a916-2a6d8186b23f · inbound

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning cites this paper.

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.896724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:50:51.653571Z digest=sha256:d38a0086fc8f55f031e606772a9e9a0621d7fcf1c3cb686ce8fd42639ec23ed4

Observation 751150a7-56e0-4b66-bdbd-90071755782a · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.655050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:63ef70f1f9c057a3672ae7d78b18ab609dc1adde6c961e13ea12da7cdca9d574

Observation 0b4e3dca-8494-4ee4-a417-c80a3138ea27 · inbound

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models cites this paper.

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:55:03.845502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T10:54:57.732141Z digest=sha256:716673799c3e6441ed8d8d424cce5b506f9113b73983eff06d852a853691e42d

Observation d03a8337-83c4-4fc6-99fc-a10f2e1ca7cb · inbound

Refining Compositional Diffusion for Reliable Long-Horizon Planning cites this paper.

Refining Compositional Diffusion for Reliable Long-Horizon Planning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:11:18.920112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:47:58.141126Z digest=sha256:5e049dca9adaeedb2f230ebf47f2e1fefef023566e568b392a4b3a64e6a51f63

Observation 160a7291-e759-4891-abc5-4bf75cc0c274 · inbound

JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning cites this paper.

JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:37:51.906678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T19:37:19.404335Z digest=sha256:4e5dccd1ffb9477d1a40d065a1c9f6083f6d79baffce91e6ec28d649a7cd15ed

Observation a00fd679-7e6d-4d2a-ac68-b370845d1fa2 · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:51.690341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T19:26:19.032362Z digest=sha256:b045b556782677cec7bba496ea9294cb3580dbf6e86e11150b1cf93370e649ba

Observation 5dd19e58-12e4-4496-a756-b80bf09c29d3 · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.674696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:35:14.348849Z digest=sha256:667f4c14fedda91dabbb33a59bc4467301d701f0b84058daac41fd9ebf537d66

Observation 2f04b9ec-b5c1-47d0-bfe7-c1b9dbe2fb5f · inbound

Scaling by Diversified Experience for Vision-Language-Action Models cites this paper.

Scaling by Diversified Experience for Vision-Language-Action Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.703257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T17:12:50.216492Z digest=sha256:bbe39a47513484cdaa224825c2f056e92bb97c6afc9698e60f397e43dc5f9bc2

Observation 52822f58-c9ec-4ed4-8971-7b239830740f · inbound

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience cites this paper.

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.658025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:47:22.504175Z digest=sha256:01d9dea597d450337d6e3a8c92f2027765f0a861dc86d37e25fc1c4b5558be5d

Observation bf1732f0-4678-40d3-8402-69278e232546 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.902987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:86ecd7abfa6c20b1e9220bfcb90b06b3f1c3507be598a01212bb66633f34e473

Observation 9b5bc7f2-3940-4368-a58f-6b6c11111dbd · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.715045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:15727794e56898de2741674f247ab7c2955b2446465dcd6658ada567a0e6dadf

Observation ee17dbea-1e62-424d-b08f-aa196c8abe32 · inbound

Reversal Q-Learning cites this paper.

Reversal Q-Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:38:49.860828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T02:30:24.951689Z digest=sha256:79965b90eb252465cd61f2f4402929c31727cba8bc35ef3e7df532c7d2af4fe1

Observation 423458cd-c8b6-4838-9d59-e3ce23867c09 · inbound

Robot Self-Improvement via Human-Video Dynamics Models cites this paper.

Robot Self-Improvement via Human-Video Dynamics Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.562830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T14:40:13.855741Z digest=sha256:c2ba2ca86e27b2d63b0df312ebb498844fdf723d6da6fc585a25da5e799be3d7

Observation 4df7f6ef-7196-4271-bc01-16b0a71e1671 · inbound

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning cites this paper.

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.123745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:19:49.473294Z digest=sha256:2eae94ce1a106f7afa1809c525d9bbb0b817a4fff7457a6531a0e2232d3d8602

Observation ac60c372-34e8-4bca-9091-38a5ec11690e · inbound

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies cites this paper.

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.533700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T09:45:55.963036Z digest=sha256:e5036aeaf80673f7f5b958007233b7de00434ede757fc3fde9e23b875152a53a

Observation da2aa04d-e3cb-46c6-a4ee-83dbef6125dd · inbound

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning cites this paper.

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.474411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:17:48.731091Z digest=sha256:8c050a86c522799b2a88b9de1ab0c294bcf18064199007d3d438de56f516ed9c

Observation f9f5b974-9bec-4850-9588-93545f270202 · inbound

Controllable Sim Agents with Behavior Latents cites this paper.

Controllable Sim Agents with Behavior Latents Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:01.976930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-03T10:41:09.285743Z digest=sha256:235d9bf10e6abac4ef53ad55d951823712b1f51002bafed017425f0a893358f3

Observation ac80c2fb-d7eb-4c50-935b-3f66ea3d10e2 · inbound

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training cites this paper.

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T06:41:57.276146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:41:57.276146Z digest=sha256:2be1545deac140e48c32088d80567521e822dee52ffe54056362f9778f3ba482

Observation e2819e81-cf32-47a0-9c56-7434a032f4f6 · inbound

VINE: Taming Generative Control Policies for Reinforcement Learning cites this paper.

VINE: Taming Generative Control Policies for Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T12:17:04.321971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:17:04.321971Z digest=sha256:05886e477ee9bf4ca89c8f9da6f4ff4081cca7d2a6c896fe11f4203841784fe1

Observation f08ee7e7-79ca-4e0b-ac81-af9759d37e18 · inbound

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy cites this paper.

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:27.374609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:27.374609Z digest=sha256:9b6548ee02cd89c82ca54a16728769257996c0ce93de73c3fbd6a2fc9f6d5972

Observation 9f3b3865-8852-4d00-93fa-510cccdb0c58 · inbound

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning cites this paper.

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:23:58.895544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:23:58.895544Z digest=sha256:dc680cf38eac33b52185bea332689ceb2184aa78149a3879950b4c816160a1a0

Observation 651e4e18-a3ac-4014-b83e-788d04d0a12a · inbound

SDDBMs: Soft Denoising Diffusion Bridge Models cites this paper.

SDDBMs: Soft Denoising Diffusion Bridge Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:38:31.344854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:38:31.344854Z digest=sha256:129c737c625ba1dc5cbba445984b0fc01da2bcf9cd6abad6efc07ebf2c30b307