Pith. sign in

Paper Citation Record · LEDGER

Exploratory Diffusion Model for Unsupervised Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2502.07279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07279 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:21:18.474336Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 717804df-2c64-4226-83eb-f028019cddf1 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.552117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.114277Z digest=sha256:18aaa49f4491f4e2b82641170e512fe6c7b15d0d266fead24b1b827c65bba82a

Observation 1186c289-b8cd-485a-8f62-13bd7c4d93da · outbound

This paper cites Tenenbaum, Tommi S.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Tenenbaum, Tommi S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.120405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.120405Z digest=sha256:68b4c58623ffa98bde99775b8c0348182da4be7034ca790b142f59db8f5abb3b

Observation 52013a13-f6fb-4638-a35d-f2950327fed6 · outbound

This paper cites Diffusion for World Modeling: Visual Details Matter in Atari.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion for World Modeling: Visual Details Matter in Atari

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.125931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.125931Z digest=sha256:5e77198759f6e09611866f63fe56bdf876a1bd635503f2ca47314eb93020c507

Observation 1c04bd1b-5546-4893-bbf7-dd6bb3d3540b · outbound

This paper cites Random polytopes, convex bodies, and approximation.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Random polytopes, convex bodies, and approximation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.526532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.132825Z digest=sha256:0d5e4db8f967acaaae931d71549e27f0ff4b6b6304ceceeb326280a6184583d6

Observation b5f258a2-ebde-445b-a142-3e292e93ebe1 · outbound

This paper cites Constrained Ensemble Exploration for Unsupervised Skill Discovery.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Constrained Ensemble Exploration for Unsupervised Skill Discovery

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:21:18.939754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.138225Z digest=sha256:9f3824d9041de1c47207a8c5a49f48b86a9563fceba70084163aaf47b7a61bbe

Observation 8f7e1583-b62b-47d7-81c5-d76e8d9c4e00 · outbound

This paper cites Exploration by random network distillation.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Exploration by random network distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.144431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.144431Z digest=sha256:fe830348b2abad945656cba537c0bec4d81a6ab088fc26f36160f488d0ea8c94

Observation 9ce0c823-a0ac-4384-888a-ad82f96ec597 · outbound

This paper cites Explore, discover and learn: Unsupervised discovery of state-covering skills.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Explore, discover and learn: Unsupervised discovery of state-covering skills

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.500380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.150828Z digest=sha256:a9032f4417aabf26b2e93f324a886a066448fc8aaed45a377bfc63faeb66c8a1

Observation 4f41f3bf-524c-4d7d-b5ec-4ce13f26c927 · outbound

This paper cites DIME:Diffusion-Based Maximum Entropy Reinforcement Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning DIME:Diffusion-Based Maximum Entropy Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.155791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.155791Z digest=sha256:4b308a16bb3f5def35a5c74c5ab9bb25c2d264aad331dfef00811b4b0946c275

Observation 749b2cf6-799b-4bcd-ad89-0bd4e41b92a5 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.160941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.160941Z digest=sha256:20117d4ed22663de59ec59ee21a14563c77979f120d41d125243dc1815965a93

Observation 193876ed-7802-427e-94db-c299f5ddd8df · outbound

This paper cites Simple Hierarchical Planning with Diffusion.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Simple Hierarchical Planning with Diffusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.166804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.166804Z digest=sha256:29a86bd2347bfac40d29d881a968c620b0d951c5230ba816fd26d8fe844e9669

Observation 249abdf1-5bc1-4cb2-b12b-91895894bbe9 · outbound

This paper cites Offline reinforcement learning via high-fidelity generative behavior modeling.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Offline reinforcement learning via high-fidelity generative behavior modeling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.484341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.172641Z digest=sha256:40a18c59eb5cc58ad30425369e381f2d7690df50e6b7f91705d29d8fdc6e27a8

Observation b00d0d5a-5bcb-4ec0-81b1-cf47f5dbac27 · outbound

This paper cites Aligning Diffusion Behaviors with Q-functions for Efficient Continuous Control.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Aligning Diffusion Behaviors with Q-functions for Efficient Continuous Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.177946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.177946Z digest=sha256:8b56e946090fb8193318d9a761bbd56b63a92a0023e084809579706e11fca65b

Observation 20988d13-2fcd-4b7a-9bad-507db77e9192 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion policy: Visuomotor policy learning via action diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.183415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.183415Z digest=sha256:71cb221ac3c8570178812e4a1b1a50b8d2e2314d070ca096cf7dc516b29bbefe

Observation b8dca37e-e625-40fc-98d4-dffeb92cb3b0 · outbound

This paper cites Diffusion Posterior Sampling for General Noisy Inverse Problems.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion Posterior Sampling for General Noisy Inverse Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.188814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.188814Z digest=sha256:8a7c269226172813426b0d65fa53fe09bf5fe8a9a098350e8e83560eec883627

Observation f1b9aff1-ba8b-4410-81e3-dbf0334efd1a · outbound

This paper cites Diffusion models beat gans on image synthesis.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion models beat gans on image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.194341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.194341Z digest=sha256:9a0709214f381ab14a269e7fad68a0bd15c92bfbe8d6551b61337166a46aaf87

Observation 5bc01a3b-00f9-4a54-837c-210713bbf508 · outbound

This paper cites Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.199380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.199380Z digest=sha256:b500f85e85eb5d73e62bfbe2e3a3bbc8a5bb0d97064e1224f1280eff2876bba4

Observation cba315ae-fa8a-4bee-abeb-1f10bb41e019 · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diversity is all you need: Learning skills without a reward function

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.447422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.204573Z digest=sha256:f93dde5e266af9159e55725f219418f03713107612e8154ec24229e52bfd2ba1

Observation 54de6cbc-71f0-46b0-a13f-32ea678baa77 · outbound

This paper cites The information geometry of unsupervised reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning The information geometry of unsupervised reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.431559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.209501Z digest=sha256:787b55e854bd84e5551ae9f114e4e78908fa7a6e900a5314c2e33848c7eb0ede

Observation e6ded4df-fa16-481d-ac4b-24d104174cae · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Reinforcement learning with deep energy-based policies

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.416165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.214553Z digest=sha256:8578df372532748c965afe9dffce04f84b9bdaef5c1612c6677b768117d4735b

Observation 8b9f12ab-cea0-4ef4-984c-bbdd091494d1 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.400294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.219668Z digest=sha256:bce8efe82a8a05e96bc3a75f4362f79e8633d439ff12c91c8ae2308c28a82180

Observation 1803d2cc-0e49-4e91-8a24-4828acd0e416 · outbound

This paper cites IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.224350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.224350Z digest=sha256:20658d7dd821d6a665f340062a71391b624ee6816c253474c277cf844ec7a0d9

Observation d90fbec2-cc75-4562-a0d7-519a7e761017 · outbound

This paper cites Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.384186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.229805Z digest=sha256:ad65eb9713e3195074dd72f823d94f635bb19b4b230ae425521374494736e099

Observation 65d6ffed-8431-4bc1-a236-6b8da9905ac9 · outbound

This paper cites Denoising diffusion probabilistic models.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Denoising diffusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.234847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.234847Z digest=sha256:a72f428f3ab80566537d8ff98782d5269cddf0f598266b8982213a0dcaf76764

Observation 09ba4fbd-d15c-448d-86d4-586560b38a47 · outbound

This paper cites Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Langevin Soft Actor-Critic: Efficient Exploration through Uncertainty-Driven Critic Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.239564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.239564Z digest=sha256:82af8afb2e3109bf3e0e0c444d239fb4478cf77f7cd8e9d4c5b0a8d6ab3950e0

Observation 64ea5b7d-a4d1-46a6-b3d7-bc072f27085f · outbound

This paper cites Planning with diffusion for flexible behavior synthesis.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Planning with diffusion for flexible behavior synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.358510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.245569Z digest=sha256:20d9edeae3e45766ff98d228432f83b49dd4433c48e7943866c4c6494b087cdd

Observation 0c81c733-6952-4b95-b4ef-2d8cb7a0c569 · outbound

This paper cites Efficient diffusion policies for offline reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Efficient diffusion policies for offline reinforcement learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.250301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.250301Z digest=sha256:09a5ecfb2bf38af54431a947c58b1278c04498c17cd37b16d4b859766354001c

Observation 3204f4ed-570d-404d-a18a-3d0ca8e4e2a1 · outbound

This paper cites Unsupervised skill discovery with bottleneck option learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Unsupervised skill discovery with bottleneck option learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.333288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.255235Z digest=sha256:269cda2fee2b2970d5fdbe97fe54880f6556abd5b55eb90efe55088a81450a3a

Observation edc71b78-a83a-4795-9c54-6c4638cda4b5 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Offline reinforcement learning with implicit q-learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.259982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.259982Z digest=sha256:ece0727c24382602bf5758e3dc95f0d5d69bf3d2e913402991a4cbb82107dd3b

Observation 73b8f274-5633-4882-8931-202c37755d65 · outbound

This paper cites Unsuper- vised reinforcement learning with contrastive intrinsic control.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Unsuper- vised reinforcement learning with contrastive intrinsic control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.306391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.264755Z digest=sha256:0eae7677d8b056a03fc5eddca32a59204c3caf22f24ff735f9133fc6bada5d43

Observation cf259c37-eb98-47a3-ad23-3cb9c6052bb5 · outbound

This paper cites Urlb: Unsupervised reinforcement learning benchmark.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Urlb: Unsupervised reinforcement learning benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.289871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.269581Z digest=sha256:4bb9c6cae91abbbc73ffa999c91cec87ef1237b92e9b3cedc909554a06e0cba5

Observation b4104a14-fb61-4278-a931-3d2289529bc0 · outbound

This paper cites Efficient Exploration via State Marginal Matching.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Efficient Exploration via State Marginal Matching

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.274465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.274465Z digest=sha256:6c602ad3ca882a55c23d549a7bdabba1e63f1d13e71884d9687fd4d10f816e36

Observation 5c091b6d-b0e9-499a-b58b-8156c8a0d9ec · outbound

This paper cites Hierarchical diffusion for offline decision making.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Hierarchical diffusion for offline decision making

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.272069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.279537Z digest=sha256:98d91ae01adb3aa281b53fd45be1bd378f4b095e1e94e5eafee295878bbccde7

Observation 5d47d839-47e1-4e91-8bf6-b317946c851b · outbound

This paper cites Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.284794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.284794Z digest=sha256:80da4abe4daddc796ebfa6dacfa138ae3700db02022592ef350b96cb0182c8b1

Observation 6c2cff4f-45cc-4e34-84b6-a9c51d4e44ad · outbound

This paper cites Adaptdiffuser: Diffusion models as adaptive self-evolving planners.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Adaptdiffuser: Diffusion models as adaptive self-evolving planners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.256958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.290217Z digest=sha256:b17af3ef1ede71e23e257abd450176618b147d4c436c453b7f6a16cceb2c6e80

Observation bb26dcbb-b4dd-4744-8451-210dee826867 · outbound

This paper cites Continuous control with deep reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Continuous control with deep reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.295604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.295604Z digest=sha256:723ee176b9097db0bef46938ae17568e8f6a15f5680995ced347e5da9735d701

Observation acacfe4a-278d-49ad-9b0e-3b4e0ab72907 · outbound

This paper cites Aps: Active pretraining with successor features.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Aps: Active pretraining with successor features

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.241608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.301156Z digest=sha256:4935ce1a214dbf42eb8fbc9171a81a1de61d2470540459725e8694f96205448c

Observation 6e3ab296-d29d-45aa-8575-4f2845740779 · outbound

This paper cites Behavior from the void: Unsupervised active pre-training.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Behavior from the void: Unsupervised active pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.226391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.307110Z digest=sha256:24bb2092b25bfdcb06a789d9c4b260d04b733c6ce6661b9fe1e1ab476ada59e6

Observation 436e4d5e-93f0-458f-9091-e6ccdef8eb7a · outbound

This paper cites Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.313075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.313075Z digest=sha256:82767495d6e740b839907858eb6d84a222af27836c940fcd3179bb69f130b4fb

Observation 43d3f187-783d-4d0c-bf6d-39429be652df · outbound

This paper cites Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.210145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.318224Z digest=sha256:fc044fdab5181e2e51e20a7deb6f8210115690e772bc6fdd9231e36f3fecc035

Observation d6dd9ac4-0c13-4e8e-ba70-a58255ea0506 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.323013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.323013Z digest=sha256:2b3505be4e3e40e48cd700cec1522ae5869ef8a175cf8abebd7bf2531a1960f6

Observation cab513be-a52c-4aa6-841e-378ec63bd8f9 · outbound

This paper cites Synthetic experience replay.Advances in Neural Information Processing Systems, 36, 2024.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Synthetic experience replay.Advances in Neural Information Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.182147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.328698Z digest=sha256:a3ec49d4a7313f78c2b382c5ebd8005ae9937215bc301436da9aa059b61fb584

Observation e1fcfedc-9fa5-464d-813c-f463941d62e6 · outbound

This paper cites Efficient Online Reinforcement Learning for Diffusion Policy.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Efficient Online Reinforcement Learning for Diffusion Policy

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.333364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.333364Z digest=sha256:0813cae3607719cca3df1d2f042a200d9ded4e5d735a8a8bf9896302a58379aa

Observation 4c5de78b-2d75-439a-83d5-9af270e71cc1 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.339718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.339718Z digest=sha256:dc1d23f0e79f7d641e19b6d092b03952a4e707cde38fa8ae38dc04d94d851aaf

Observation ad88db33-c4b8-4834-933b-e230cbf6614d · outbound

This paper cites Curiosity-driven exploration via latent bayesian surprise.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Curiosity-driven exploration via latent bayesian surprise

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.166694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.344637Z digest=sha256:072ed126cd1a71752cc95122777b763e4fa016eb66d1d1e0b722455f523c14e5

Observation e36c7829-ae27-43f0-8d40-9e40197bdf24 · outbound

This paper cites Lipschitz-constrained unsupervised skill discovery.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Lipschitz-constrained unsupervised skill discovery

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.151235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.349181Z digest=sha256:90283fdec447c2f3499d05a1a62a625b2a8b09baa2891f47044fce122e14cd0b

Observation 25988992-dc49-4b2d-b2c8-54d202233e5e · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.353849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.353849Z digest=sha256:1d2493e86196dc6306af4aba94c6a3a4a240627b037c981284afbf3351884dc4

Observation de555d34-5ec6-4044-b0ab-c4e20ecf47c2 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Curiosity-driven exploration by self-supervised prediction

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.135131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.358862Z digest=sha256:2a00f9f7fc814f98514aa9887daf1caa8df678baea0fa85ba86b6b3fe5395951

Observation 360a91dc-9430-4d88-ad63-b63cf128d0fd · outbound

This paper cites Self-supervised exploration via disagreement.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Self-supervised exploration via disagreement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.118151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.363809Z digest=sha256:94e108e99ba1ebe668be7507d87d1444ebb5388e3f98f4f25dd943dd79212c6b

Observation b957b3af-87ce-4b1b-b3ef-1e63cc384755 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.368484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.368484Z digest=sha256:c707c2d06912ae52f94c9209dd4f4140375afc3e9ed8b4c3d181c208bc1b0711

Observation 42042865-f0ce-403b-a6f5-260f47240610 · outbound

This paper cites Learning a Diffusion Model Policy from Rewards via Q-Score Matching.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Learning a Diffusion Model Policy from Rewards via Q-Score Matching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.374992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.374992Z digest=sha256:982be10db22b3ad242a693ccb7eb36daff676dfe4dbccad619df1563089ee733

Observation 06931fde-8a40-45aa-87d0-7d1de4e1d46c · outbound

This paper cites Diffusion Policy Policy Optimization.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion Policy Policy Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.380367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.380367Z digest=sha256:3ae852f19c48846be3f0187336f71e8f14b3a0d30d75c4b8f4424a9adc307230

Observation 2707d75e-063d-4aec-be49-31b33b7b274d · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Photorealistic text-to- image diffusion models with deep language understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.385361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.385361Z digest=sha256:40c5e3b42818fbaa6484fb5ed8c1fff9cc8eb5140d5de8e75bf12c3f457a61a1

Observation db3205d2-3abc-442f-97a3-48c031397268 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.390208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.390208Z digest=sha256:037f7f942b8ad4e1dc5171d275066566325d37829a97d79797d3c87a7fb3d343

Observation cbbf8218-1a96-4e75-b11d-10589a6f930c · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Deep unsupervised learning using nonequilibrium thermodynamics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.395721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.395721Z digest=sha256:43603b3286424356f59b43729d08f7287fe818befc199b2d1a24aeb3faca3327

Observation 9c1c4d90-8598-49f0-8f25-d4866102963a · outbound

This paper cites Denoising diffusion implicit models.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Denoising diffusion implicit models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.400244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.400244Z digest=sha256:a4409d872c5b140e78e7d5885e14aca459d9c183afaa6d61ef0554048360139e

Observation 413bb471-9086-4afb-904a-9a6070a93c95 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Score-based generative modeling through stochastic differential equations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.404894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.404894Z digest=sha256:058439b2bc10dc1224feb7dc1451d3708cbff71aed262d2ddf3a4e9d475faac7

Observation 0b3437f2-eb5b-4c53-89d8-2a520ae2f4b7 · outbound

This paper cites Reinforcement learning: An introduction.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Reinforcement learning: An introduction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.409791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.409791Z digest=sha256:3cc8bc4e478cf79675615ecf15cc6f570ea0e7c00059fb66f345be4204916a35

Observation 55ee6be8-6836-40f9-a899-c5f22e3d4788 · outbound

This paper cites DeepMind Control Suite.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning DeepMind Control Suite

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.414534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.414534Z digest=sha256:a6bee1646bfd6a3a9f0ba17776527212c6ddc1643deba37c69e934ed339da5dd

Observation bce0a378-af0f-4c70-a587-9312f03ea012 · outbound

This paper cites Prioritized Generative Replay.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Prioritized Generative Replay

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.420119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.420119Z digest=sha256:f83eb63c833147f8a318a23f5f583d33aed1244faecfdcda4326d69e08b1f662

Observation 6594c9c2-0ac2-47fb-9fb0-2a70427a63b8 · outbound

This paper cites Diffusion policies as an expressive policy class for offline reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion policies as an expressive policy class for offline reinforcement learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.424969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.424969Z digest=sha256:d3f58a0acf489a9503553688426b1ad686b03d1454cbe3d7c54afcbb693eb801

Observation 0198e1c6-8411-45ae-951a-24aca56b6b25 · outbound

This paper cites A problem in geometric probability.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning A problem in geometric probability

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.039086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.429537Z digest=sha256:108a9d383b5c7b6f28862c8c5722cc4166685c99018d6de960d43774e14338f3

Observation 5944f1a6-f8ea-44a9-90e7-e9a2249fdadb · outbound

This paper cites Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.434308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.434308Z digest=sha256:939092e904ba7103f8030e5192d38c258ad83a1cca3e6d15f515aa3e513a548e

Observation 3f30d08a-8ca5-447b-b349-8d05edb0bc5b · outbound

This paper cites Policy Representation via Diffusion Probability Model for Reinforcement Learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Policy Representation via Diffusion Probability Model for Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.438947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.438947Z digest=sha256:4aaf94a1b0b6243076a0ca58a2882232166ec05b6aa622c0b1013d43b5268292

Observation 78c1c545-d060-40a5-be76-0c7b9837b1b8 · outbound

This paper cites Behavior contrastive learning for unsupervised skill discovery.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Behavior contrastive learning for unsupervised skill discovery

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.022950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.443810Z digest=sha256:387d8a18bb2b9c9978b953e636c6d93f52c078b39136d0e1af3fdac774b6baf0

Observation a5ec26a0-6739-42d7-9d04-108188dd472f · outbound

This paper cites Peac: Unsupervised pre-training for cross-embodiment reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Peac: Unsupervised pre-training for cross-embodiment reinforcement learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:19.006929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.448752Z digest=sha256:1a4eeeb39d9c49440ab1f91c1b271257409055c8d645d431763401fd234cbc8a

Observation 6ba8d291-d21f-4541-8de4-667a1bc26e33 · outbound

This paper cites Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Towards Safe Reinforcement Learning via Constraining Conditional Value-at-Risk

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.453588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.453588Z digest=sha256:01fa8bd6956dd846c1ec18ff0f870677a210195408179bc20072734bbf81be94

Observation 5376eb95-3c8f-4f54-84ab-3ed8f27c2477 · outbound

This paper cites Automatic intrinsic reward shaping for exploration in deep reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Automatic intrinsic reward shaping for exploration in deep reinforcement learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:18.989867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.459069Z digest=sha256:7edece2e13ccdf1639b011a92c6bb94cf59ce831b0a73c708000de1e2aeeac48

Observation 97a9345c-5642-4da6-94fc-5352cf84b381 · outbound

This paper cites EUCLID: Towards Efficient Unsupervised Reinforcement Learning with Multi-choice Dynamics Model.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning EUCLID: Towards Efficient Unsupervised Reinforcement Learning with Multi-choice Dynamics Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T13:21:18.464264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.464264Z digest=sha256:9f94328033f403cf2b2d8b8a4376a08baf42a5b6584bec3590deea1482c1d91a

Observation 7d7c98b1-be79-47c8-b37e-c988112b88fa · outbound

This paper cites A mixture of surprises for unsupervised reinforcement learning.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning A mixture of surprises for unsupervised reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:21:18.973702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T13:21:18.469375Z digest=sha256:cd62f27b82eaef2bfe5cc55731cad656563a04cfc1a4e3cf356792cac2700469

Observation a33c2e8f-595f-448c-ba50-fecddc257fad · outbound

This paper cites Diffusion Models for Reinforcement Learning: A Survey.

Exploratory Diffusion Model for Unsupervised Reinforcement Learning Diffusion Models for Reinforcement Learning: A Survey

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-08T13:21:18.474336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:21:18.474336Z digest=sha256:8ec1869131bfd41e73d1c70b5c4b47791545b070765e8919a637f51c0a9f8f23

Pith citing papers

No inbound Pith citation observations are available.