Pith. sign in

Paper Citation Record · LEDGER

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization

As of 8 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00795.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00795 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:26.101533Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac9d45cc-7a9f-41d9-b013-0c0be567dee7 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep reinforcement learning at the edge of the statistical precipice

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.717787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.717787Z digest=sha256:eab7187c9f8b0388f70fb10f5d066ef74ae2ce2c2a49697a5be1cc8564bbbd38

Observation 0c629fa2-b676-4ebf-a403-7d3f06ebde26 · outbound

This paper cites On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.723026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.723026Z digest=sha256:b2c0f9621d7be327ce7ffc0be64a5e4a77e9922ccdc3328d62768e24d7981d74

Observation a6022746-c35d-493e-ab30-e9d86c9fab73 · outbound

This paper cites Hindsight experience replay.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hindsight experience replay

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.728070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.728070Z digest=sha256:5640c28e5416ee2b98bc1d2751d44dbff8405f3d2f778056ea9f6712b3a26a06

Observation 0eec7250-8f73-4841-8163-0584315efc57 · outbound

This paper cites Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.733026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.733026Z digest=sha256:1176da05750f5f36368c9fddab2e22c02c41b1588d9ab5d73c815c7eac53a6dd

Observation d0775145-8232-416d-ad3c-985501eafa1c · outbound

This paper cites Accelerating goal-conditioned reinforcement learning algorithms and research.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Accelerating goal-conditioned reinforcement learning algorithms and research

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.123474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.737574Z digest=sha256:a6c5499386c67c115ee18e16dc3be765a8558fc593d8f53d123dbb11134a564d

Observation fb9f4169-e126-4d26-8760-662962c2a265 · outbound

This paper cites Offline rl without off-policy evaluation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline rl without off-policy evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.742340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.742340Z digest=sha256:525027987dbfd73512c1c8c97460a476ee0bcf9bca88c670f9fd29085ab93c6e

Observation dc72a2a9-7eda-4695-8aeb-59ff29e0fce6 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.100466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.747851Z digest=sha256:f723ad2b3604ec14d255ac238f48aaa2e0810de000b2fda9a9cc743671e38f93

Observation 1900dc52-3d95-4889-bea5-aa5ba811b9a9 · outbound

This paper cites Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.752674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.752674Z digest=sha256:de83d24b5974bab7917ae1d85cf443baf54c86dd1d890e6eb5a1760ec32fdd69

Observation da09bc93-76e4-4615-9e8e-9c8fe5094a62 · outbound

This paper cites Goal-conditioned reinforcement learning with imagined subgoals.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned reinforcement learning with imagined subgoals

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.088998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.757238Z digest=sha256:1043e192c45213d43193ba5d8b6b9e649a335257b75ee5429f5b1ba393e0f705

Observation d46ca468-7388-4700-9cb4-81b704157246 · outbound

This paper cites On the statistical benefits of temporal difference learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the statistical benefits of temporal difference learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.073974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.761572Z digest=sha256:5d1d7ac057b15c7adcac0468836f92a166fcc59324d3f8eec30b315cb31dad9f

Observation 05f055f6-a886-4d96-9ed4-26e1783a705a · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision transformer: Reinforcement learning via sequence modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.766013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.766013Z digest=sha256:80656df5a1cbba0fd966c6bddb47b0b351019fd98de92262f21ae3aba5c62183

Observation 40ee4e50-4a6d-48ab-a17f-d624539b2950 · outbound

This paper cites Goal-conditioned imitation learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned imitation learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.770153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.770153Z digest=sha256:6c9199da1274023420293ad3808548368f098cc5168e830c86a9e9ebbbac24c4

Observation b26c610a-459b-44af-9acb-69a25dc55670 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization RvS: What is Essential for Offline RL via Supervised Learning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.775369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.775369Z digest=sha256:1f7102869bd7b47264b32209a663f67cd52d0c6b8122c1ee48f25b6257150bb2

Observation bbad15d8-706f-4444-9661-dcec1aa5946d · outbound

This paper cites C-Learning: Learning to Achieve Goals via Recursive Classification.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization C-Learning: Learning to Achieve Goals via Recursive Classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.785189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.785189Z digest=sha256:3e5564a2e56618b7ed0940c7f5959e4affa0ad408fa7fdc7e2fca21c2fdcaf55

Observation 6ffaf888-2f16-4a8a-8ffd-9d2ed56273b1 · outbound

This paper cites Imitating past successes can be very suboptimal.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating past successes can be very suboptimal

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.041974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.797884Z digest=sha256:88323dc2ded51f46e61eb79a6a716653a35ff55e014dc4ea061c1bcdbfb0a214

Observation 1b94d514-e345-483e-ab0f-72493bc2707a · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive learning as goal-conditioned reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.026589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.806254Z digest=sha256:6f43615563c0bbef17affa995ee813bbe19982de03aa9a17533bcc9c0ea0d917

Observation 8408698a-ec22-445f-8d72-8d4e4f62cfc2 · outbound

This paper cites Inference via interpolation: Contrastive representations provably enable planning and inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Inference via interpolation: Contrastive representations provably enable planning and inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:27.009713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.812950Z digest=sha256:963fa7dd31be962efb775af5273fdbf2a795e461d498e5f4dc94f002e4fd4e7d

Observation 4e9db73b-ece7-43c6-9a15-e5a3a1b6028b · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.820597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.820597Z digest=sha256:5610f167527f9133de914ae6706a0dff838927366bd9ef3fddcc5cf521311d39

Observation 6833a941-aa87-4b87-b910-42a1c54cfb00 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A minimalist approach to offline reinforcement learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.828824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.828824Z digest=sha256:8e022aa21687f7caeec147b24f3120a15fba57cc45b94ff94530119405b37bf1

Observation 68403ab2-45bf-4103-8b10-8bdc7ef7f3bc · outbound

This paper cites Learning to reach goals via iterated supervised learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via iterated supervised learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.986664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.836914Z digest=sha256:bf631246ee1b6a21959f6021d0a68340295ae0e588edc9765ab8e2c4a4963281

Observation 2e8b731c-3160-440d-9f66-c4cc18ed1276 · outbound

This paper cites Closing the gap between TD learning and supervised learning - a generalisation point of view.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Closing the gap between TD learning and supervised learning - a generalisation point of view

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.973047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.845241Z digest=sha256:97619a4f96bfab637a69d469aec10432f476ced37acbc79400d90f7fb25ac1d7

Observation 36e41fa6-edff-4f22-af06-65760ef38fe7 · outbound

This paper cites Distance Weighted Supervised Learning for Offline Interaction Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Distance Weighted Supervised Learning for Offline Interaction Data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.851471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.851471Z digest=sha256:ade28756770413665c382ad97d761a5e434f4a7d9fb6cbfe2311346dacbfba6d

Observation f7e355dd-8db1-498a-bbed-21c4b1f7bf24 · outbound

This paper cites Diffused task-agnostic milestone planner.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Diffused task-agnostic milestone planner

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.960312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.856909Z digest=sha256:5f4236eee0edda5fcbea20d74337aa3e255e0ed586f6d142c5c19a40fe62a9c4

Observation 77b97c48-a143-4bf1-abcd-26486fc6b117 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:05:26.394932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.860929Z digest=sha256:bbbaf660c5a6176d240ca3477c0e44ba4801b3e05cb3915c9ca9feb5a3193d99

Observation d2b47e03-40b7-4d65-91e9-1887d2c8aa79 · outbound

This paper cites Learning to reach goals via diffusion.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via diffusion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.949447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.865082Z digest=sha256:52fa03e50d93b2ead14bd35c6bce83783cd00bf97341e832daa04fc0e26d9881

Observation 95e9d632-ee99-466d-abfc-0287ee48a294 · outbound

This paper cites Efficient planning in a compact latent action space.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Efficient planning in a compact latent action space

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.938587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.869691Z digest=sha256:a024ef98a1cbaeac5d0c39a6e1b5d2d31bf87b18921e50c081bd76da71c83416

Observation 333012e5-a28a-43fa-9fbd-f51f78732e31 · outbound

This paper cites Adaptive q -aid for conditional supervised learning in offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adaptive q -aid for conditional supervised learning in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.927879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.874197Z digest=sha256:9d69633915cfe596dea0fa67001a024d0753fbb2caec42da8f2b8b8b1fdca54d

Observation 7cb756c9-6bfd-4cd5-993b-d1ddc942d258 · outbound

This paper cites Imitating Graph-Based Planning with Goal-Conditioned Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating Graph-Based Planning with Goal-Conditioned Policies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.878623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.878623Z digest=sha256:4c0b91383fb37f83b0a7d28165c5b2f5d1415378a940365b6365c8758c7a9ac5

Observation 991722db-63c3-41eb-83da-9416adc5569c · outbound

This paper cites Adam: A method for stochastic optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adam: A method for stochastic optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.916192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.883186Z digest=sha256:c9cb5235daacfcedbec3e674dc93a148021d83e05bd102d673330d6cce6cee20

Observation 3f22aaa2-4fed-44ab-a087-56d7c8e1c94a · outbound

This paper cites An Introduction to Variational Autoencoders.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization An Introduction to Variational Autoencoders

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.887182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.887182Z digest=sha256:b9df9daf57cbd96d28ed2b9ebd699f3abe6426ba903965ad54b744c90964f002

Observation 0d1a09a9-e890-4949-a7ad-ce801bfbbe15 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning with Implicit Q-Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.890730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.890730Z digest=sha256:2e5d0a9716f09865c6c6c74c0a1d8aa73e33f15d9bb6eda168453fd2a4b2d08e

Observation 6cac748e-1609-4fa2-8267-547de17c0893 · outbound

This paper cites Stabilizing off-policy q-learning via bootstrapping error reduction.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing off-policy q-learning via bootstrapping error reduction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.894463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.894463Z digest=sha256:8e4b0f9e0efcc70ce20abd9cc2da7c79fef8c7d3de4566a6282e86da319e370e

Observation e129b63b-db0d-4810-b0ba-659543a77553 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Conservative q-learning for offline reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.898096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.898096Z digest=sha256:cca8ecb1b0a04197f666a1dfdcf2d34d18a0533e92cddc25b9660c67bad100b9

Observation 298c4dee-8b0d-474e-afad-c57e481b4223 · outbound

This paper cites Multi-game decision transformers.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Multi-game decision transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.890075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.901393Z digest=sha256:c28b93c709fff05cf6f927ac398e339bc623e7e99f17d676ba495c6ab292f182

Observation 4da2b345-5e0d-4d95-a039-aed07cae0242 · outbound

This paper cites Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.877877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.904989Z digest=sha256:1a02a02bc2923abb96de277966b11624fbafe3f971c69e8596d2d417cf993688

Observation d0173034-cc60-47ef-acb7-cf86d44b445c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.909047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.909047Z digest=sha256:aaa35d0cee808a3956eb70dcb4e51d967eaa38b4cff0a0c182c7ef2574235468

Observation 338ffa32-a5ea-4dcf-8a0d-aae0f181c06a · outbound

This paper cites Hierarchical planning through goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hierarchical planning through goal-conditioned offline reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.865153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.912606Z digest=sha256:eeb29ce5b97fd91f9d45388e43952561401e5946f4405c6dce36714efe9e5635

Observation 2b3f252c-1f37-403b-ba3f-b2186e6ef988 · outbound

This paper cites DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.916444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.916444Z digest=sha256:76c1ace18457af619539446f9208e99ea3794b5d1d30cd1446a8267f8e4b9873

Observation ab0ce511-34ba-4e69-b886-0908976a8819 · outbound

This paper cites Beyond ood state actions: Supported cross-domain offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Beyond ood state actions: Supported cross-domain offline reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.850298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.920305Z digest=sha256:05b9b0dccbcef0a8281d07898e56da6e2292b6a0b0baa01a913346031840ed2b

Observation d1625f51-008d-47f1-a75b-80e9ec423a6f · outbound

This paper cites Least squares quantization in pcm.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Least squares quantization in pcm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.924152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.924152Z digest=sha256:80b96897742a4c102e0bce66131bd19c004fb15653c7de4cd534c91f82f7102f

Observation 05f6f010-fa26-43cc-9226-12825383213b · outbound

This paper cites Decoupled Weight Decay Regularization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decoupled Weight Decay Regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.928161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.928161Z digest=sha256:e59b67c11050888d2fd609dcffc52183b789a877a9d0f78ecb989ed369db1da6

Observation 2a9bff04-ecfd-42a1-be87-47952d4c0d07 · outbound

This paper cites Generative Trajectory Stitching through Diffusion Composition.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Generative Trajectory Stitching through Diffusion Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.931670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.931670Z digest=sha256:0cb70cccb34717d760053e4c0ba1b87ad1ec74388e4b493aa2ccd18bc6470cad

Observation 0ee3c592-d739-4461-ae0c-81d25979e657 · outbound

This paper cites Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.823899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.936092Z digest=sha256:035800d6a0a0b47742989923e66358779236674a2090932fabb044d84b1c5704

Observation 9161b203-157a-43c2-a75a-e6c58b3bf523 · outbound

This paper cites Learning latent plans from play.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning latent plans from play

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.807672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.939663Z digest=sha256:e7bbd35871963d6018c9d4868576080d0bcd79e533b634eee44896a3fd819c3f

Observation 7760b84f-ef4a-4c07-b158-dd9ed90e7346 · outbound

This paper cites How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.943342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.943342Z digest=sha256:4171d235179bbee0feb649fc1503b02778c782d356f4e67c556ed06e4abef705

Observation c121fe31-7789-4cdd-bd85-6da077e82bd2 · outbound

This paper cites VIP : Towards universal visual reward and representation via value-implicit pre-training.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization VIP : Towards universal visual reward and representation via value-implicit pre-training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.787378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.947029Z digest=sha256:70cda8f290ed20a8fe0199edbec1df251c4a9619a2ad54d8c13b3e6359689205

Observation 81535569-6821-4cf1-b589-a68adfbf95cf · outbound

This paper cites Human-level control through deep reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Human-level control through deep reinforcement learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.950640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.950640Z digest=sha256:8b05fecd96116f3dfadb281fdd6bb18932c83d46648ea98867538b5399ab1c36

Observation b836f8df-5d17-418a-9393-f6a1a8ccaa13 · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.954383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.954383Z digest=sha256:aca34d5e4bf4ed560377c5f9f15d6447b07e5d1fcf4e95166be73823b3443edf

Observation d110249c-705f-4d70-a3d4-ceec22004d39 · outbound

This paper cites Asymmetric least squares estimation and testing.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Asymmetric least squares estimation and testing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.763153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.958453Z digest=sha256:b654c1c096d2863eb599ddfc8259e178599bfdbdf9606f3bf7690d5a7e8a6f92

Observation 6fff31ef-0216-4c08-a905-3cc793992213 · outbound

This paper cites Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.961759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.961759Z digest=sha256:a05a8b43780444f9467134ac6930b392abf8788a686ce612ba7fce5d4d1b18d1

Observation 49179c0b-5541-4a55-b2de-51de3d7635b3 · outbound

This paper cites Hiql: Offline goal-conditioned rl with latent states as actions.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hiql: Offline goal-conditioned rl with latent states as actions

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.966571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.966571Z digest=sha256:a8778283c7a9357f5c617d55636be1697b956947dd05ef49da8b0adacad45c71

Observation f5f5575a-0a9d-432b-9163-2d9147e6f3df · outbound

This paper cites Foundation Policies with Hilbert Representations.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Foundation Policies with Hilbert Representations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.970149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.970149Z digest=sha256:3db2bceb36f522992842feb00d4674cd7875e5267b0ae5bd93d6228b31047b37

Observation 0b6d3909-af36-4f36-b8d2-07deef469bd8 · outbound

This paper cites Ogbench: Benchmarking offline goal-conditioned rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Ogbench: Benchmarking offline goal-conditioned rl

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.739722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.975726Z digest=sha256:0032b2a7e895c9bf19db4893dffd7265ecca563857d38948bb57031591323bff

Observation 03cde593-9e7a-4b38-b8c0-d2b09ffe2fe9 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A survey on offline reinforcement learning: Taxonomy, review, and open problems

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.727985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.980393Z digest=sha256:a3b6f9bb8f67b9025e4118ac1f195d3b0abcbcca19f7596a537da3aea537c225

Observation 738e246b-af48-488e-bd3d-3f9c6928ce3d · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.984273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.984273Z digest=sha256:f33b9daf9f8bcd6eb5d1295c5d4e39a702d2bdbf54c37e2c6851932d429b094f

Observation 0047489f-1d2b-4ace-91ba-3d824f4a0da4 · outbound

This paper cites Stochastic backpropagation and approximate inference in deep generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stochastic backpropagation and approximate inference in deep generative models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.716004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.988205Z digest=sha256:49ca79ed3160f62e5af32ce4547a26c9a6db0d7fa2f5bbbcc7fd96a42dfc4688

Observation f43cd76c-e99b-47ef-a3bc-6bd55c4b6680 · outbound

This paper cites Outcome-driven reinforcement learning via variational inference.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Outcome-driven reinforcement learning via variational inference

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.704287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.994354Z digest=sha256:74247cab61426460c3a55b7d3132c9cd4f87e9fb68b4fc6f3bd561bdc218b676

Observation 5d75398f-09e4-484f-a1a0-d8b10e8421f6 · outbound

This paper cites Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:25.998229Z digest=sha256:e6253657474406e0ea8ed154aca85ea5884ee927451a0404e99ee1e3da9ef93c

Observation a0ad97e9-02cb-4edd-8211-cb21fa4dab4d · outbound

This paper cites Rapid exploration for open-world navigation with latent goal models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rapid exploration for open-world navigation with latent goal models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.681724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.002289Z digest=sha256:e65b49e9d687bcc6c56c97e803ee6f5a571ebc314b66b82e1281b2d71f5cf609

Observation dfc8bf9e-530b-400e-a12c-6bd0d5fc4e6f · outbound

This paper cites Score models for offline goal-conditioned reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Score models for offline goal-conditioned reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.670794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.006634Z digest=sha256:b5aaaca9cc87e12fe18067124b09887f1fae488c491dc6517bff562050afe28d

Observation a9e70b2b-72cf-44f6-996f-cfd0632dad84 · outbound

This paper cites Geoadditive expectile regression.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Geoadditive expectile regression

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.659954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.010660Z digest=sha256:5a1e8fe9cdd1ea09d349d0e6ca3a850e65e5040090d2225091bc4192a4594ec8

Observation d9c00165-7d4f-4aeb-9178-49005ea0b73e · outbound

This paper cites Learning structured output representation using deep conditional generative models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning structured output representation using deep conditional generative models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.014377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.014377Z digest=sha256:525d74c60f3bce15024f44c09f684bd3d5cdefcd7ddf842b510902ba1a16e6e1

Observation 31cf407e-cc0a-4685-ae51-73d4d982f234 · outbound

This paper cites Gymnasium (mar 2023), 2023.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Gymnasium (mar 2023), 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.639864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.018409Z digest=sha256:85ff768fd47befd0bf8ca8ebbdb57a8acc30eac4150d9f44dd676e6f25290f56

Observation 15c15ca4-5ecf-46a0-abb1-078650721ce1 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep Reinforcement Learning and the Deadly Triad

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.022196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.022196Z digest=sha256:684c2ca47ee1fb869c1231c2eb066cf93c205c81c1ce1ec839ab2a3f5d8f558f

Observation 7e47b5a1-f2c6-4aec-b59c-e20c712aadcc · outbound

This paper cites Attention is all you need.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.026210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.026210Z digest=sha256:6d703669ee7a0ddf4717992189a77d7ecde6cec3f09d8d6c15a45f75915ce0b6

Observation c20b3ea7-d6cd-4f80-a058-c02026c8f25b · outbound

This paper cites GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.619349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.029722Z digest=sha256:342c35f0e7dabd7bb6710f582129a2fc6e6a2a217b73fbbaed00758e787962cd

Observation 34fa6ff0-f0f9-4b0f-bac5-f5c5f9efce2b · outbound

This paper cites Optimal goal-reaching reinforcement learning via quasimetric learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Optimal goal-reaching reinforcement learning via quasimetric learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.607776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.033672Z digest=sha256:f056290683553068b446f14316a46ff07ecb8e3d5e036518eaa85ac233a003ec

Observation 4064caf9-0460-452a-a414-d3ea921924ee · outbound

This paper cites Critic-guided decision transformer for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Critic-guided decision transformer for offline reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.596625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.038097Z digest=sha256:243701d2f5fa88c92ac2ac2c1a94454a3bf56d15d415d7b06a9f4f1cd1678550

Observation 1f240bd1-3f82-44a7-900f-bb0dcb6eb6ec · outbound

This paper cites Supported policy optimization for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Supported policy optimization for offline reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.584880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.041485Z digest=sha256:afe5d2caacd24cb8f1b73c059c8edc20ee887074666e899538ab0ee4ccb7b0ff

Observation e8e42e3d-1cdb-42c8-9228-c14463d124a2 · outbound

This paper cites Elastic Decision Transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Elastic Decision Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.045615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.045615Z digest=sha256:ef4659ad773e83b7b917fd09005e5a5b51b5706600ed9d29eeba80d2561f7ba9

Observation ca8c62b3-2b08-4809-9704-568102ecae9d · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.573852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.050121Z digest=sha256:24d856696821fe5fb3147fc66d25930a867fe341780c251b06211abad9a9eefe

Observation 158ff2de-8365-47d0-809c-dfad1cc405e2 · outbound

This paper cites Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.054551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.054551Z digest=sha256:137712b501ab947c5859ffefbf13c555a221395b1eecd3afa5335e992051efb9

Observation db0fa5c6-5581-4b25-87fb-013fcef04677 · outbound

This paper cites What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.561418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.058489Z digest=sha256:82d61aadb83318f3417b47f78987cc192a7803465e656762f0a9eb2d5e51cd6b

Observation 77db9b91-9fca-46a6-9296-56fca0933486 · outbound

This paper cites Swapped goal-conditioned offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Swapped goal-conditioned offline reinforcement learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.062222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.062222Z digest=sha256:67403e0e27f183b457fd6448c0bf553b4fd86ecceca9d853c42f23663b30caeb

Observation 2f8dd9a5-e9c1-4d76-af4a-b181000ab2fb · outbound

This paper cites Breadth-first exploration on adaptive grid for reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Breadth-first exploration on adaptive grid for reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.066117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.066117Z digest=sha256:305b669e4870088decd9463bfc39aa0508672bba22be9a882f3a1816876aa24e

Observation 7fbcd7cd-bb20-4774-be4e-c2841fc8cbee · outbound

This paper cites Goal-conditioned predictive coding for offline reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned predictive coding for offline reinforcement learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.539901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.069683Z digest=sha256:78ab9b87f66cb223adb06425601a64436e2ecbe8be6fce514fccb81db255fb68

Observation b155832e-295f-432b-bfcd-1de1323f54b8 · outbound

This paper cites Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.072975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.072975Z digest=sha256:ba5ae1608867275768c6184ebca46c46a23f5ba00bab0191c273a8c32ff24dc9

Observation 1b355547-cf41-4fd4-860b-354833c8961b · outbound

This paper cites Contrastive difference predictive coding.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive difference predictive coding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.526281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.076893Z digest=sha256:aa42f6605b9399b32bd2172aabd1fbebef5b4e1f0507683e9aac7dc684c3adff

Observation 1007d93e-a759-4670-a9ae-ce0f34545b99 · outbound

This paper cites Online decision transformer.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Online decision transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.514001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.081476Z digest=sha256:cea16d0844043b8d191fd3e6b05bdc682868e4793f23b71253afe198b2b4dbdb

Observation 2de92e6a-ec70-4159-9a24-cf51af5f8e47 · outbound

This paper cites Behavior Proximal Policy Optimization.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Behavior Proximal Policy Optimization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.085023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.085023Z digest=sha256:859bc6a10d5895e4af06bea80e970eb053ff98f836389cc405668e76513d62c3

Observation 1267e842-1074-4441-b705-b7277e3483cd · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.091332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.091332Z digest=sha256:8644a0366fe9597e4751664feaf8f85710ca946c12566643e914300e41d0deda

Observation ad6054d5-8274-4dd4-b327-6f158fb0db69 · outbound

This paper cites Revisiting the design choices in max-return sequence modeling, 2025.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Revisiting the design choices in max-return sequence modeling, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:26.500024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:05:26.095894Z digest=sha256:4c33402054c6c872ddd39978ae26c58caf312a404bf9163f97f117d719dcf345

Observation fc578017-cce3-4fd5-bb09-46cadeecab2d · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Maximum entropy inverse reinforcement learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:26.101533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:26.101533Z digest=sha256:cad3ef09dc88ebf06ce6277cb5b6e55a816022fba0db4147160829410b8b7931

Pith citing papers

No inbound Pith citation observations are available.