Pith. sign in

Paper Citation Record · LEDGER

Diffusion Guidance Is a Controllable Policy Improvement Operator

As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 34 inbound Pith citation observations for arXiv:2505.23458.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23458 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:30.265947Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:42:22.223628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.122193Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 274acc93-49ec-4872-82e3-7fc79ccf1802 · outbound

This paper cites Diffusion policies for out-of-distribution generalization in offline reinforcement learning.IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies for out-of-distribution generalization in offline reinforcement learning.IEEE Robotics and Automation Letters (RA-L), 9:3116–3123, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.230721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:24.913040Z digest=sha256:cc2bb9ba98ca582d7fc9e1f519621652e36c209f12a2d1abd936aff15cb8e736

Observation 0e25982c-fdbb-4d15-a09c-5f11796d28ca · outbound

This paper cites Building normalizing flows with stochastic interpolants.

Diffusion Guidance Is a Controllable Policy Improvement Operator Building normalizing flows with stochastic interpolants

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.216123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:24.978487Z digest=sha256:73fc2c5503a34d20299b8173511167ded8daa7abadcf66a18a8aead7cc9b05e9

Observation 8c8a5edd-a8ff-4ab6-bfc9-6f2bd24dac69 · outbound

This paper cites Uncertainty-based offline reinforcement learning with diversified q-ensemble.

Diffusion Guidance Is a Controllable Policy Improvement Operator Uncertainty-based offline reinforcement learning with diversified q-ensemble

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.201252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.066955Z digest=sha256:7a20977f70cd9d71806589939bc5d7233ddf55c8b42c8e33fa599e67dd8e6d7a

Observation d631dc44-c4c3-4791-a139-0e2f163d7972 · outbound

This paper cites Hindsight experience replay.

Diffusion Guidance Is a Controllable Policy Improvement Operator Hindsight experience replay

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.186004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.131856Z digest=sha256:5435a3ebb489b848e7904c5a2ba4a8a00b18859772555d43c0d95687ea3cc00f

Observation 9d261bb7-52e2-42b6-9e15-7cd904e73583 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Diffusion Guidance Is a Controllable Policy Improvement Operator $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.208106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.208106Z digest=sha256:e239500b8fd3220ebf181707151a1d8b03487f2a2e4831db8a300985f02dc6b6

Observation 58bcf97e-f0ce-4eb1-80b2-b42e8b204c9b · outbound

This paper cites Zero-shot robotic manipulation with pretrained image-editing diffusion models.

Diffusion Guidance Is a Controllable Policy Improvement Operator Zero-shot robotic manipulation with pretrained image-editing diffusion models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.296803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.296803Z digest=sha256:303544ea38e69e35cb337fbeb5e7cb1544edd6d9fdf5bf85c857be7fa1347d86

Observation 714c8e21-4f4a-42e1-88a9-710d62b6844c · outbound

This paper cites Whitney, Rajesh Ranganath, and Joan Bruna.

Diffusion Guidance Is a Controllable Policy Improvement Operator Whitney, Rajesh Ranganath, and Joan Bruna

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.162096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.405048Z digest=sha256:9556cf2f757ea8e631f8ce9a71e00405e94c374f0858b8b4914a2899c9aa42dc

Observation c8d9265a-1c4c-48e5-9356-e219f334da01 · outbound

This paper cites Offline reinforcement learning via high-fidelity generative behavior modeling.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning via high-fidelity generative behavior modeling

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.145715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.509248Z digest=sha256:5b9b6d9c887c8ae1bd83586a6e7a6e1d75cf3854fc4718c1920598d9fb0bce7a

Observation 74bde7e8-c041-4700-92e2-603a020bc4e2 · outbound

This paper cites Score regularized policy optimization through diffusion behavior.

Diffusion Guidance Is a Controllable Policy Improvement Operator Score regularized policy optimization through diffusion behavior

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.131431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.589564Z digest=sha256:5bddec50c1b3aea2caf12e5f0042dc3e14ebf9f2cdaae70d7b7dc4937de3815f

Observation 802a466e-2092-4a69-b502-dbd8eef22ece · outbound

This paper cites Aligning diffusion behaviors with q-functions for efficient continuous control.

Diffusion Guidance Is a Controllable Policy Improvement Operator Aligning diffusion behaviors with q-functions for efficient continuous control

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.115183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.704932Z digest=sha256:374a7b473f29172fff43d6b6002a03d5df22f041bff938b27d825c841a3cd1ae

Observation 31dbe662-30bc-4e13-a288-4e83934ed286 · outbound

This paper cites Abbeel, A.

Diffusion Guidance Is a Controllable Policy Improvement Operator Abbeel, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.099654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.778783Z digest=sha256:0381c060863296f82933c954c9679c9211a39bafd17edbd928db5ee8eaa48ffe

Observation 2e5c6a4c-cd75-47ac-a79d-c4ac3e03c3fc · outbound

This paper cites Diffusion policies creating a trust region for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies creating a trust region for offline reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.084718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:25.891954Z digest=sha256:a99229489b4bbedf2a871e585668395a3f62f1df0c82ea00da8edcdc5f3d82e8

Observation f99072af-9eb1-449b-aa3c-70647f1bfc58 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy: Visuomotor policy learning via action diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:25.998739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:25.998739Z digest=sha256:9f33917648393c3da308498c2ee628ef6c73752f38539f05facf72baf5d4e9be

Observation 642ed821-922b-4b5d-a86d-7d811cc1397f · outbound

This paper cites da Silva.

Diffusion Guidance Is a Controllable Policy Improvement Operator da Silva

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.059739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.110964Z digest=sha256:7221769da250e4a71fc8e464161c5efec31d1b6ce5d9753a45e12595f5d0f11f

Observation 6660750a-a6d4-4d59-a499-5734479665b2 · outbound

This paper cites Using expectation-maximization for reinforcement learning.Neural Computation, 9:271–278, 1997.

Diffusion Guidance Is a Controllable Policy Improvement Operator Using expectation-maximization for reinforcement learning.Neural Computation, 9:271–278, 1997

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.044191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.195008Z digest=sha256:9cbd30a3b3b422e347e4116214fcfbec5ddc107b955303f428c341338f842793

Observation 6baffa69-bd7b-42ef-9032-355ef9e9716a · outbound

This paper cites Diffusion-based reinforcement learning via q-weighted variational policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-based reinforcement learning via q-weighted variational policy optimization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.028196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.273415Z digest=sha256:51bf8ec00e50c39797481664863a6f5bd07eeff8ea6cb62a4647a2e013c21816

Observation d2ae4034-f701-45e1-b67f-ce913f910fe1 · outbound

This paper cites Consistency models as a rich and efficient policy class for reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Consistency models as a rich and efficient policy class for reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:32.012294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.379402Z digest=sha256:a31b4f65cfcbb55c6ddcd8eab431f1275b41c7bc3b90381b851a2ec37a5ca850

Observation 3dfa2dba-d65f-4a2f-b469-1a53c1577de1 · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations (ICLR), 2022.

Diffusion Guidance Is a Controllable Policy Improvement Operator Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations (ICLR), 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.994233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.455926Z digest=sha256:dbb0d97c80af5acaa0890d64840a67798e7fc9a3be94f4ac313a2faacbbf412a

Observation c051a217-5a4e-4391-aacb-46e6b944f61f · outbound

This paper cites Imitating past successes can be very suboptimal.

Diffusion Guidance Is a Controllable Policy Improvement Operator Imitating past successes can be very suboptimal

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.977730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.558212Z digest=sha256:2d917aa2c4b2b0b14b514142f75d4de2a6d5a9d3aea805a27c727ed3ea99b501

Observation 5256e09e-c9a8-4162-a405-424367c63127 · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive learning as goal-conditioned reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.961567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.631969Z digest=sha256:631bb911e786f9d6a856f5249d0de64b2f09ad3900d21e7d7688dbab41c28726

Observation ab7d341c-fbe0-4bb5-bcbe-8476cdff4e70 · outbound

This paper cites Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion actor-critic: Formulating constrained policy iteration as diffusion noise regression for offline reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.945632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.695872Z digest=sha256:07908795a6b8e6f89b6a761809afcab6f66808d0bbed7f0f68b00d736d462dfa

Observation 69bb8e44-52f8-4eaf-8a65-28adba2ab38f · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator A minimalist approach to offline reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.930897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.785311Z digest=sha256:5cf6cf717fe484eccfc9286698d69c7669e7a6de68cc6887cadf36fa7d88f1e1

Observation 8d237293-255e-4cc9-bb01-6abf487dfdd2 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Diffusion Guidance Is a Controllable Policy Improvement Operator Addressing function approximation error in actor-critic methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:26.868584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:26.868584Z digest=sha256:f0d994cf46ec4a95642930d5158139382d08310c122cc95667ba1c94e06101e7

Observation a683e37b-e065-4cb8-a17c-da924fc5e915 · outbound

This paper cites Murphy, and Tim Salimans.

Diffusion Guidance Is a Controllable Policy Improvement Operator Murphy, and Tim Salimans

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.906617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:26.952030Z digest=sha256:531f49e72f8853b353648dfa08ac1db06caae389b70813924d3b7898ad12d9a2

Observation 00d8e094-27f7-4188-ae32-abe1b12d04a9 · outbound

This paper cites Extreme q-learning: Maxent rl without entropy.

Diffusion Guidance Is a Controllable Policy Improvement Operator Extreme q-learning: Maxent rl without entropy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.891327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.012832Z digest=sha256:b7a7bb2f2b6dfde9b84b0c8d382fe7284ce6c59c5549aee43d5b0fd445f495dc

Observation 5a58fd3f-0340-408a-aa74-34bfd3a72932 · outbound

This paper cites Learning to reach goals via iterated supervised learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning to reach goals via iterated supervised learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.875520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.077763Z digest=sha256:2a9cb868d85813810cdd9ea8a97e99a9216c82456110b062cd4fa3afaeb2767a

Observation f936e271-bb74-451b-9a60-626cda99cec8 · outbound

This paper cites Closing the gap between td learning and supervised learning–a generalisation point of view.

Diffusion Guidance Is a Controllable Policy Improvement Operator Closing the gap between td learning and supervised learning–a generalisation point of view

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.860415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.159496Z digest=sha256:d2965fca5c5a4f1c9cb9cc854833a2b93663b88cfbd0c2a4cc461f10748e5df0

Observation b890cd14-cc24-450f-8c7a-cb04512583e1 · outbound

This paper cites Explaining and harnessing adversarial examples.

Diffusion Guidance Is a Controllable Policy Improvement Operator Explaining and harnessing adversarial examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.219561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.219561Z digest=sha256:fba0b038c2553bd21640f5358ade238f438c8dc245064940ea9a19cb1cfed33a

Observation 0af6ac9e-9dba-4833-81fe-212a580d3a1a · outbound

This paper cites Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.835362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.306393Z digest=sha256:004d94bab8bf6a9c8e7b710ec5bf43fa5eebde2e13ba724a2ac45d50bbc500bb

Observation a609deab-73a5-4a8b-8378-70e300f433dc · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Diffusion Guidance Is a Controllable Policy Improvement Operator IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.377097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.377097Z digest=sha256:720d81daf57f748d35d17397805782d18b12fcd94abe591c114498363a04a726

Observation d80827a7-1c30-4b8b-b0e5-1a18224f6796 · outbound

This paper cites DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator DiffCPS: Diffusion Model based Constrained Policy Search for Offline Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.429381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.429381Z digest=sha256:111ab6bf503b236e058039cdbcad5266cc91210a4f9a71b109564ec72c50e45a

Observation 5b45d15d-891e-4072-840a-4d8ea3a004fd · outbound

This paper cites Aligniql: Policy alignment in implicit q-learning through constrained optimization.ArXiv, abs/2405.18187, 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Aligniql: Policy alignment in implicit q-learning through constrained optimization.ArXiv, abs/2405.18187, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.509049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.509049Z digest=sha256:dbaf99d872e60d6b1fcd00f1c313cdc9cc12331bd286def12ccbe5356541778b

Observation 4cdc18df-65ac-4108-9745-750ac70b7148 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Diffusion Guidance Is a Controllable Policy Improvement Operator Gaussian Error Linear Units (GELUs)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.600669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.600669Z digest=sha256:f8538727ae5388aeae75c9bd5979f53a9ff74797c12d65a25f2f02b785d5fde4

Observation 53feb01d-6d0a-48da-8d91-cd9763944c6d · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diffusion Guidance Is a Controllable Policy Improvement Operator Classifier-Free Diffusion Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:27.693635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:27.693635Z digest=sha256:1fd47f0b3b410ac014e6b51134feda399e65567f48b815b805a3992fcf52fb60

Observation 2b83f3ae-92e2-49cb-86b3-f8b5cf6b2d18 · outbound

This paper cites Denoising diffusion probabilistic models.

Diffusion Guidance Is a Controllable Policy Improvement Operator Denoising diffusion probabilistic models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.818742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.756726Z digest=sha256:576a0b4842cc0e4fb158cfcd13368085e2ccd7f0f187e99155f6b43d2f6fd76a

Observation 79375b08-a493-42d7-b055-4c6e4e25eb82 · outbound

This paper cites Tenenbaum, and Sergey Levine.

Diffusion Guidance Is a Controllable Policy Improvement Operator Tenenbaum, and Sergey Levine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.803927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.830267Z digest=sha256:b37124e0b78fffb9b360bd9423da81e9fb07de194d50e68103386fe1d06c464c

Observation 04dfc25a-5712-416c-ad4c-bb68502e31b6 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Efficient diffusion policies for offline reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.790011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:27.927436Z digest=sha256:82936e933f64f3cfcc17a38dfb215296e00575a2b6f8eec39923d82f992f1387

Observation b0a23421-8d9a-4410-a78f-015dc9ea282a · outbound

This paper cites Kingma and Jimmy Ba.

Diffusion Guidance Is a Controllable Policy Improvement Operator Kingma and Jimmy Ba

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.003819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.003819Z digest=sha256:c45a89fd4da79f36b9bbcb8319917ef287407e5e5b4ead77013f9a70e2e3fae8

Observation 1f1d601c-2551-4ac9-8bc4-695c4a5526b3 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline reinforcement learning with implicit q-learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.766144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.120329Z digest=sha256:99cac36a552d3828def6a1e1e10681a52b5ade5529103be7d781de63f4943d72

Observation 19b2c0ff-aedb-4e00-8e0e-01f2ac6f42ab · outbound

This paper cites Advantage-conditioned diffusion: Offline rl via generalization.OpenReview, 2023.

Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-conditioned diffusion: Offline rl via generalization.OpenReview, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.751229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.181310Z digest=sha256:455a511decf50dd2863f406ee8fb785e66bb3f15ea86a82fe04caaea33dbec0f

Observation b68cdd13-872c-4b13-9aac-37ecd00af636 · outbound

This paper cites Reward-Conditioned Policies.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reward-Conditioned Policies

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.274947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.274947Z digest=sha256:080f106d08beab349fa322f626bcbf8fb9be7a3abf1a84f122f5d741a687a202

Observation 8ab444d0-758f-4913-84b3-886758f587d9 · outbound

This paper cites Tucker, and Sergey Levine.

Diffusion Guidance Is a Controllable Policy Improvement Operator Tucker, and Sergey Levine

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.736331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.365220Z digest=sha256:426ab5d9862c4c7a7d47c8c28f7eee5b9549dd3abfb9cd9db256d9e3a233fe46

Observation 54fb33ee-57a4-4ae5-ba20-ccd08657ac2d · outbound

This paper cites Learning multimodal behaviors from scratch with diffusion policy gradient.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning multimodal behaviors from scratch with diffusion policy gradient

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.720162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.446855Z digest=sha256:c9da600388620506790ea88084d4321a83d19b2469496f3bcbf212f7334202ff

Observation f14ddcfc-602d-4fa8-96ae-14e7d2bddbfb · outbound

This paper cites Lillicrap, Jonathan J.

Diffusion Guidance Is a Controllable Policy Improvement Operator Lillicrap, Jonathan J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.705213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.512443Z digest=sha256:5314e3a7a19731c9adb223cfeff7ab7eb734a844f31caffa19e5043daec79625

Observation c7778eff-81cd-448d-adc2-3761d12556a8 · outbound

This paper cites Flow matching for generative modeling.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow matching for generative modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.536332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.536332Z digest=sha256:a0499f45a7d12884c2331e4c5d7b963073c065b77a7788b0eeac6f712a4701b2

Observation 40be38c3-64eb-46d2-be15-b5110bc091d2 · outbound

This paper cites Flow Matching Guide and Code.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow Matching Guide and Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.549452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.549452Z digest=sha256:6e90e101f589bac98bae7e6f9b9af56d08fbffae67a3f43eec76a05000c30de4

Observation c15fd0fc-e658-40a5-9f74-445ad8470372 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.554370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.554370Z digest=sha256:7d5672ab24f989da18c5378b473fd5e411c09bac54172eec39e763c41b95fd3c

Observation eb4ac65c-1a4a-45fd-89c7-228952e56585 · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.672148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.574898Z digest=sha256:dfdab77df6c6ae1781dfa368d8cdc3c63f9ba6158a144921f23e6e3996b39ae1

Observation 186a8a5e-bffe-49d5-8301-714d5d616d02 · outbound

This paper cites Learning latent plans from play.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning latent plans from play

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.653436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.653436Z digest=sha256:17381d6ecfce06381376639b2dd688f3a48600a3b6b51ae1a241248746dc7266

Observation 8b39392f-0dca-4393-9244-8e326a2cd847 · outbound

This paper cites Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.647886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.713236Z digest=sha256:3791df2e3fe801a692f014be1bca558747c1c8367850f61f427f2a30ea63f28d

Observation a7725c5d-50d7-4222-a48e-3c5cdce9f246 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.792421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.792421Z digest=sha256:942aace3b6cff0848d7d98e098a9cc7f9240543aaea5fed759857300759624a8

Observation ad88bfd8-f5c4-412d-b5d6-4f3426b19802 · outbound

This paper cites Mish: A self regularized non-monotonic activation function.

Diffusion Guidance Is a Controllable Policy Improvement Operator Mish: A self regularized non-monotonic activation function

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.633448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:28.875496Z digest=sha256:4c8647b0478bf603f154a917ce799b56ad88bfad75bcb9f0bd5b3b219813a482

Observation dc9e6f52-cd4f-4fc8-b454-080ea9473496 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Diffusion Guidance Is a Controllable Policy Improvement Operator AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.945562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.945562Z digest=sha256:23329c580095fb6ff3f925fbece58fae739088e9ff97f4d006bc76a01a19f52e

Observation 63304ea1-0b86-4062-ae8a-489c57bf65e0 · outbound

This paper cites Anti-exploration by random network distillation.

Diffusion Guidance Is a Controllable Policy Improvement Operator Anti-exploration by random network distillation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.618102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.026219Z digest=sha256:eb1928bf3444746fbde6fdc44d102c8e37f627e2f695f02a362a98e350852d9b

Observation aff38423-441a-48b5-a61c-d9a8530723e8 · outbound

This paper cites Is value learning really the main bottleneck in offline rl? InNeural Information Processing Systems (NeurIPS), 2024.

Diffusion Guidance Is a Controllable Policy Improvement Operator Is value learning really the main bottleneck in offline rl? InNeural Information Processing Systems (NeurIPS), 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.602591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.036298Z digest=sha256:a5659c06ce734efc19d53e1ae3bb473ba4b0e5292bcc1cdf8e739c76fb5006bf

Observation e8f0d878-ac34-4dda-bbd3-57460468c85f · outbound

This paper cites Ogbench: Benchmarking offline goal-conditioned rl.

Diffusion Guidance Is a Controllable Policy Improvement Operator Ogbench: Benchmarking offline goal-conditioned rl

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.040779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.040779Z digest=sha256:ae677e332045620a6118ebbdbf2c278cc91b885dd7b270a58c25c8b33ddb05b2

Observation 47c62e4a-7ec3-413d-b363-d7a933bb5ccd · outbound

This paper cites Flow q-learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Flow q-learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.577999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.065263Z digest=sha256:24b0c75ca15a0dc7b30b36bb74f98fb4f9017e6a523b470ec2718aff3843ca48

Observation 7c802dd2-e7f0-4e3a-a7c6-b8d2d3ce315d · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.148020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.148020Z digest=sha256:dd735323a04949ccecca6da35ac9a22761e33f3d671b44cb2fd808c0973ef3ad

Observation 2a6fe649-1620-49cb-9503-9196a9bba391 · outbound

This paper cites Reinforcement learning by reward-weighted regression for operational space control.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reinforcement learning by reward-weighted regression for operational space control

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.562100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.220323Z digest=sha256:225fd90d829ce137a5f324c239eb12d76b3304420ecab255287b588af4f44cd1

Observation ea3880fc-0749-47e5-bbe3-6c9b3a1ad795 · outbound

This paper cites Learning a diffusion model policy from rewards via q-score matching.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning a diffusion model policy from rewards via q-score matching

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.547890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.278182Z digest=sha256:e42dc7915282c9b9a01d1d2dd35de026e25dce46e03ca06699e772e545d47f52

Observation 31f9eb45-0ce3-4031-be9e-a5a4fe774f09 · outbound

This paper cites Diffusion policy policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policy policy optimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.534055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.356106Z digest=sha256:69554c657acbf732b00d6e7576870712d94418626e10647de03084b603b0544a

Observation ddb9c88b-0d86-4d36-9e6d-b96d693575d6 · outbound

This paper cites Trust region policy optimization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Trust region policy optimization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.520084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.437576Z digest=sha256:812e844243d76b46f129839c9be88b49665ca305421d5586781c6badd02b2251

Observation 5794becc-5d9e-4908-a0c2-dc1fb1b1ee4d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Diffusion Guidance Is a Controllable Policy Improvement Operator Proximal Policy Optimization Algorithms

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.495573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.495573Z digest=sha256:6f123799f372d860bd7a2fedb7a4f0eebe54069b3d01795989cc4ed8a5fa6a30

Observation f63e00e8-fe76-4f37-977e-baef64f090e5 · outbound

This paper cites Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum.

Diffusion Guidance Is a Controllable Policy Improvement Operator Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.506492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.534334Z digest=sha256:bc55f8dd24cce0203ada8326b75ff7d363f34428bc6e289d4c10e14e199cccfd

Observation d8c70b07-4674-4b80-ab56-1f72b7ca04a8 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Diffusion Guidance Is a Controllable Policy Improvement Operator Deep unsupervised learning using nonequilibrium thermodynamics

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.538385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.538385Z digest=sha256:c6cd929acfcc63d8b85f7ce7f63916fd6d9013602bda352f0e49a03144e754f8

Observation ae689195-a547-4cfa-a0ef-fcc6fba41660 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Diffusion Guidance Is a Controllable Policy Improvement Operator Generative modeling by estimating gradients of the data distribution

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.483749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.544585Z digest=sha256:c61842931424c544fd3a2e9bf36b4e01d2ee68da0bf7b582b439adf1548260f0

Observation 5871684f-043c-42a1-a6d1-2b1e394593d7 · outbound

This paper cites Sutton and Andrew G.

Diffusion Guidance Is a Controllable Policy Improvement Operator Sutton and Andrew G

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.469657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.612569Z digest=sha256:674b8d453d69c4f3f784888177f629b9fd3becd41a70278c659f3758757f3bb3

Observation 6b9578f0-d2a8-466d-8172-62933aaaca77 · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy gradient methods for reinforcement learning with function approximation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.442908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.689334Z digest=sha256:35978797d0e2200e91437a40d1a3321cb07f13b548c8830f438ae8f1c446ca30

Observation e82727e7-5032-43ee-9501-5bc00cdd03b7 · outbound

This paper cites Revisiting the minimalist approach to offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Revisiting the minimalist approach to offline reinforcement learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.760477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.760477Z digest=sha256:8aad6c626fefaadc0b57f4095231eaaa502a5d85c142638d50a6b3f780abf299

Observation 8204b323-f5b4-4fe3-a956-2fbf79d60459 · outbound

This paper cites Learning one representation to optimize all rewards.

Diffusion Guidance Is a Controllable Policy Improvement Operator Learning one representation to optimize all rewards

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.328623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.799586Z digest=sha256:935c6c1ca945c166c36675fb0060a45a1823ff2e0bd670c79cd7be7cd8fc484d

Observation 44ea1e02-a513-4732-a729-ed5ce1f0510c · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.182390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.858233Z digest=sha256:fc64ccbc8feedb9ac35fcd8bcbc9cad00b643d696bfdd8a074dbd4e6562287e6

Observation 594f885e-0130-465b-a3d4-70f16a80e98d · outbound

This paper cites Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Manfred Otto Heess, and Nando de Freitas

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:31.106882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:29.933259Z digest=sha256:2f2a9c03e6b16f9d9c3f977812174ac9f97c69a4d99b60168d7be20e0d2e0e14

Observation e7adf455-ba68-41b8-ab50-8952b9fb234c · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Behavior Regularized Offline Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:29.993903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:29.993903Z digest=sha256:b039b2a5f0e6aeeb9bd963bcc0471a59e06ce94bb51b0c1258f318d0c4e36e40

Observation afcc1f24-a588-4efc-ba1a-105e0309a826 · outbound

This paper cites Offline rl with no ood actions: In-sample learning via implicit value regularization.

Diffusion Guidance Is a Controllable Policy Improvement Operator Offline rl with no ood actions: In-sample learning via implicit value regularization

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.997347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:30.055766Z digest=sha256:2914f9a2971075ea4fff575161d53f801cadd857826faeb29ef2bf4b023258f8

Observation 481580db-6b44-4321-8133-98cbe73ed38e · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Diffusion Guidance Is a Controllable Policy Improvement Operator Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.876244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:30.064184Z digest=sha256:cf84174923e9d73761d099eb9dc5667fc5b0110e37e0b46764b7d7d9fae287da

Observation 749c8f5a-2bf8-4433-b336-ba5affe8f977 · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:30.069061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:30.069061Z digest=sha256:d3c6bb095d4caba0ecada128f19546ac4cf3685cea21831eecea75f3f6458560

Observation 062be4fa-d7d1-4753-afa9-7000e4ee3e37 · outbound

This paper cites Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Don't Change the Algorithm, Change the Data: Exploratory Data for Offline Reinforcement Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:30.096727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:30.096727Z digest=sha256:9e3542468e876984398c377834e8227ca4cbc4b004d289ce01591581eee9014d

Observation b7214780-13b5-4915-a75a-6de01e8e131c · outbound

This paper cites Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.703344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:30.159233Z digest=sha256:9a4c7b1a8c37f4ba302ef3fa387eb30458ca5a4d4e256019b6a7d5454bad8391

Observation fd15f4a3-d5a5-42db-ba9d-5a09e4d6c2c5 · outbound

This paper cites Energy-weighted flow matching for offline reinforce- ment learning.

Diffusion Guidance Is a Controllable Policy Improvement Operator Energy-weighted flow matching for offline reinforce- ment learning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:30.673893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:56:30.265947Z digest=sha256:3579032890c61f52e8706fe3671f013c96a1f0dee18d380d709d1d2c80fc1513

Pith citing papers

Observation 56e107ec-6e8c-4c85-a1b2-0b52646da785 · inbound

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance cites this paper.

Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:22.223628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:22.223628Z digest=sha256:01bb8fc789d2a0e29cc974e393a7cd3ba441f0115b07d7db53d51d36c33358cd

Observation e09793a3-be45-46dd-ab6f-0509b118ac7c · inbound

DiffusionNFT: Online Diffusion Reinforcement with Forward Process cites this paper.

DiffusionNFT: Online Diffusion Reinforcement with Forward Process Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:54:30.998401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:54:30.953199Z digest=sha256:c5d28d13e6da590234571733adfb51afb9d0a8aaea41e22f5ac23b99a1410114

Observation c674f213-37de-4d0d-bacb-bec5bad07779 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.250729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:d15c3a44b080b0c81fd7292c19d0b25d5cdc6bb501ed79bc57c9361b0ef3555f

Observation 21705f32-f9ce-49ff-a85b-0490e878683c · inbound

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator cites this paper.

Latent Reasoning in TRMs is Secretly a Policy Improvement Operator Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:29.374560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:29.374560Z digest=sha256:9e237228708d6a82e4e7d36d126721c29e1fe5e418e74289873ed2a73c47aeee

Observation cba2875a-4da8-40c8-be8f-2d6167778d44 · inbound

Dichotomous Diffusion Policy Optimization cites this paper.

Dichotomous Diffusion Policy Optimization Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T13:18:29.364170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:18:29.364170Z digest=sha256:48b19afaa17e64b9990ce87e3768b16ab7abe350869907f2ce4d5eb557c3498e

Observation 1b108cbb-24e0-4cb4-a360-d36d9a30a11f · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:30:31.938553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:ab85c52cec876b568f1b0aa44f9f1e5efbcdf101f32e3a1e97b5af4ebfeea4c0

Observation c50286fe-3ad8-49ce-8090-0f3f49962bf5 · inbound

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training cites this paper.

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:47:54.752480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:47:54.752480Z digest=sha256:c61e64341348dc4b24281180b03de8a21767ab8368341c2d7955a70fb3d3c56a

Observation 6feff226-b30e-4dc5-ad74-e425c919bf0f · inbound

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control cites this paper.

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:06:43.044814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:05:39.797848Z digest=sha256:510fd8c9fb60cdc42871e3785fe96d065a51a13d6fc03c14f64fbb40e2ba38c6

Observation 9dae9f6f-7c72-481c-868b-5f8291ec48e8 · inbound

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning cites this paper.

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T23:47:45.866615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:47:45.866615Z digest=sha256:ded7bafade75b0fb4b8bd805364ecaec808433e6a8c7d2f4a2aa643875db0552

Observation 54c53f7d-15fd-40a5-b96b-befcb10abac6 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:374efc7850fc34fe5dc5067885ef072fbe599cc09d71f46fb36dd78d73936c6f

Observation 99916ee1-1754-44d0-ac18-3fa2ad6c56d5 · inbound

Update-Free On-Policy Steering via Verifiers cites this paper.

Update-Free On-Policy Steering via Verifiers Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T23:44:37.937519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:44:37.937519Z digest=sha256:e7cd5f4c0b338a19eeb473eb6814f072e97936fb420692807871e040cbf93a36

Observation cc3c10a8-86bf-4b35-95a9-20bf1f2df124 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:25:59.160153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:7f956d0d92887e509bf889391e55818240adb107ef0937e9a5efcc4969b15e6b

Observation b7a92b01-d2b5-4d81-a199-b009fea7e531 · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:817babacc8d40beb1a8f935f2a65a367395b694dee30ac6041d61aa586ae8250

Observation 116dcf5c-f634-4dfc-a916-2a6d8186b23f · inbound

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning cites this paper.

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:05:58.896724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:50:51.653571Z digest=sha256:3328574c644196f3b36a7603f1b74382783ccddd6d104db6ddb579f74015454c

Observation 751150a7-56e0-4b66-bdbd-90071755782a · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.655050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:6e0a73109e2d0830da5f1e967961ccd3be231cdf00fa8de304dace791557f615

Observation 0b4e3dca-8494-4ee4-a417-c80a3138ea27 · inbound

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models cites this paper.

Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:55:03.845502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:54:57.732141Z digest=sha256:671c4a5be7a8687fc206d2ef8f9fff2def88751da829e8e207eddd19914a7c89

Observation d03a8337-83c4-4fc6-99fc-a10f2e1ca7cb · inbound

Refining Compositional Diffusion for Reliable Long-Horizon Planning cites this paper.

Refining Compositional Diffusion for Reliable Long-Horizon Planning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:11:18.920112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:47:58.141126Z digest=sha256:074b533907f4933cb78ddca0624fdf0896e806de637fc8315f5cd776db8b8e3f

Observation 160a7291-e759-4891-abc5-4bf75cc0c274 · inbound

JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning cites this paper.

JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:37:51.906678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:37:19.404335Z digest=sha256:0fa62f197b30f89287eed1170151b73e39b14131be4bd9e6410cdf34af297f28

Observation a00fd679-7e6d-4d2a-ac68-b370845d1fa2 · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:27:51.690341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:26:19.032362Z digest=sha256:1448b55845d0de9c83bde5e918f98686ec8aa4f838b5f6d2bae47ae4fb52da93

Observation 5dd19e58-12e4-4496-a756-b80bf09c29d3 · inbound

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy cites this paper.

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.674696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:35:14.348849Z digest=sha256:3d8b974066831edfb2e7d3222a265525442d4bd5344eed5b87d0be1b71754e30

Observation 2f04b9ec-b5c1-47d0-bfe7-c1b9dbe2fb5f · inbound

Scaling by Diversified Experience for Vision-Language-Action Models cites this paper.

Scaling by Diversified Experience for Vision-Language-Action Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.703257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:12:50.216492Z digest=sha256:11befae22934e761f01b70724400ba87cc972f305caf90fce63f0b5af5e621fd

Observation 52822f58-c9ec-4ed4-8971-7b239830740f · inbound

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience cites this paper.

DexPIE: Stable Dexterous Policy Improvement from Real-World Experience Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.658025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:47:22.504175Z digest=sha256:9ad7db7a576839198c1cc65dcd53758d0658cf8eaae98e52ab52eb9c31da493e

Observation bf1732f0-4678-40d3-8402-69278e232546 · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.902987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:4db6fc44d2cebf486eaaef46ba302d6263d4115c0be53e779a200b9708b3c543

Observation 9b5bc7f2-3940-4368-a58f-6b6c11111dbd · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.715045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:7a3b5aa154181067cea2369ca9f770efc4d36879fad62ba03a865a300e299027

Observation ee17dbea-1e62-424d-b08f-aa196c8abe32 · inbound

Reversal Q-Learning cites this paper.

Reversal Q-Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:38:49.860828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T02:30:24.951689Z digest=sha256:f6568a88c88601f67bd0489e7cc97474513079617ee5a983cb2ed86df417ce87

Observation 423458cd-c8b6-4838-9d59-e3ce23867c09 · inbound

Robot Self-Improvement via Human-Video Dynamics Models cites this paper.

Robot Self-Improvement via Human-Video Dynamics Models Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:37.562830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:40:13.855741Z digest=sha256:6da26c231f05cae8e98871746d7d479979259e22808c40fe43ad2b7e436c2e3b

Observation 4df7f6ef-7196-4271-bc01-16b0a71e1671 · inbound

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning cites this paper.

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.123745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:19:49.473294Z digest=sha256:9f37d7007b9868851b911e71725c1601b18252d8a923e8f1b8bd988d1e6a3759

Observation ac60c372-34e8-4bca-9091-38a5ec11690e · inbound

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies cites this paper.

ReGuide: From Test-Time Guidance to Self-Improving Diffusion Policies Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:54:40.533700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:45:55.963036Z digest=sha256:b5d127956825be424f17b535e6501f07b0233da13558f45d9f5cf3af12bf89a6

Observation da2aa04d-e3cb-46c6-a4ee-83dbef6125dd · inbound

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning cites this paper.

STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.474411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:17:48.731091Z digest=sha256:b3f0f11cad97effc8c82343e27b63db3ee37fb7af73aba6edde34cc994004f19

Observation f9f5b974-9bec-4850-9588-93545f270202 · inbound

Controllable Sim Agents with Behavior Latents cites this paper.

Controllable Sim Agents with Behavior Latents Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:01.976930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T10:41:09.285743Z digest=sha256:130edaff30b870ef97c8adbb2c5df30a35fde697f76955b455b8b38545954e53

Observation ac80c2fb-d7eb-4c50-935b-3f66ea3d10e2 · inbound

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training cites this paper.

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T06:41:57.276146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:41:57.276146Z digest=sha256:b5bef8e408ac82802366c80dfda624a3258929cd974bbab79951dee578cae5b8

Observation e2819e81-cf32-47a0-9c56-7434a032f4f6 · inbound

VINE: Taming Generative Control Policies for Reinforcement Learning cites this paper.

VINE: Taming Generative Control Policies for Reinforcement Learning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T12:17:04.321971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:17:04.321971Z digest=sha256:4e968ddf620699956206669125b36ba36939e4225209e41166cd40ee84dd0eed

Observation f08ee7e7-79ca-4e0b-ac81-af9759d37e18 · inbound

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy cites this paper.

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:27.374609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:27.374609Z digest=sha256:ac8a5bcc99d166484def2ff0bd4f853c8ed740fd063455ebae855341717d8cc4

Observation 9f3b3865-8852-4d00-93fa-510cccdb0c58 · inbound

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning cites this paper.

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning Diffusion Guidance Is a Controllable Policy Improvement Operator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:23:58.895544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:23:58.895544Z digest=sha256:d9ef6a1e80c76231f7aff7f0f807b769548628b3b6f20e7f8a4f91fd23880bc0