Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:26.101533Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2506.00795.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:26.101533Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
83 of 83 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ac9d45cc-7a9f-41d9-b013-0c0be567dee7 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep reinforcement learning at the edge of the statistical precipice
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c629fa2-b676-4ebf-a403-7d3f06ebde26 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the estimation of production frontiers: maximum likelihood estimation of the parameters of a discontinuous density function
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6022746-c35d-493e-ab30-e9d86c9fab73 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hindsight experience replay
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eec7250-8f73-4841-8163-0584315efc57 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0775145-8232-416d-ad3c-985501eafa1c · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Accelerating goal-conditioned reinforcement learning algorithms and research
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb9f4169-e126-4d26-8760-662962c2a265 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline rl without off-policy evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc72a2a9-7eda-4695-8aeb-59ff29e0fce6 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35: 0 1542--1553, 2022
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1900dc52-3d95-4889-bea5-aa5ba811b9a9 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da09bc93-76e4-4615-9e8e-9c8fe5094a62 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned reinforcement learning with imagined subgoals
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d46ca468-7388-4700-9cb4-81b704157246 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization On the statistical benefits of temporal difference learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05f055f6-a886-4d96-9ed4-26e1783a705a · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision transformer: Reinforcement learning via sequence modeling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ee4e50-4a6d-48ab-a17f-d624539b2950 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned imitation learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26c610a-459b-44af-9acb-69a25dc55670 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization RvS: What is Essential for Offline RL via Supervised Learning?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbad15d8-706f-4444-9661-dcec1aa5946d · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization C-Learning: Learning to Achieve Goals via Recursive Classification
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ffaf888-2f16-4a8a-8ffd-9d2ed56273b1 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating past successes can be very suboptimal
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b94d514-e345-483e-ab0f-72493bc2707a · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive learning as goal-conditioned reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8408698a-ec22-445f-8d72-8d4e4f62cfc2 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Inference via interpolation: Contrastive representations provably enable planning and inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e9db73b-ece7-43c6-9a15-e5a3a1b6028b · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6833a941-aa87-4b87-b910-42a1c54cfb00 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A minimalist approach to offline reinforcement learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68403ab2-45bf-4103-8b10-8bdc7ef7f3bc · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via iterated supervised learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e8b731c-3160-440d-9f66-c4cc18ed1276 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Closing the gap between TD learning and supervised learning - a generalisation point of view
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36e41fa6-edff-4f22-af06-65760ef38fe7 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Distance Weighted Supervised Learning for Offline Interaction Data
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e355dd-8db1-498a-bbed-21c4b1f7bf24 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Diffused task-agnostic milestone planner
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77b97c48-a143-4bf1-abcd-26486fc6b117 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2b47e03-40b7-4d65-91e9-1887d2c8aa79 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning to reach goals via diffusion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 95e9d632-ee99-466d-abfc-0287ee48a294 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Efficient planning in a compact latent action space
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 333012e5-a28a-43fa-9fbd-f51f78732e31 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adaptive q -aid for conditional supervised learning in offline reinforcement learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7cb756c9-6bfd-4cd5-993b-d1ddc942d258 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Imitating Graph-Based Planning with Goal-Conditioned Policies
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991722db-63c3-41eb-83da-9416adc5569c · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Adam: A method for stochastic optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f22aaa2-4fed-44ab-a087-56d7c8e1c94a · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization An Introduction to Variational Autoencoders
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1a09a9-e890-4949-a7ad-ce801bfbbe15 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning with Implicit Q-Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cac748e-1609-4fa2-8267-547de17c0893 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing off-policy q-learning via bootstrapping error reduction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e129b63b-db0d-4810-b0ba-659543a77553 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Conservative q-learning for offline reinforcement learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298c4dee-8b0d-474e-afad-c57e481b4223 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Multi-game decision transformers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4da2b345-5e0d-4d95-a039-aed07cae0242 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Dhrl: a graph-based approach for long-horizon and sparse hierarchical reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0173034-cc60-47ef-acb7-cf86d44b445c · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338ffa32-a5ea-4dcf-8a0d-aae0f181c06a · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hierarchical planning through goal-conditioned offline reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b3f252c-1f37-403b-ba3f-b2186e6ef988 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization DIDI: Diffusion-Guided Diversity for Offline Behavioral Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0ce511-34ba-4e69-b886-0908976a8819 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Beyond ood state actions: Supported cross-domain offline reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1625f51-008d-47f1-a75b-80e9ec423a6f · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Least squares quantization in pcm
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f6f010-fa26-43cc-9226-12825383213b · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decoupled Weight Decay Regularization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9bff04-ecfd-42a1-be87-47952d4c0d07 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Generative Trajectory Stitching through Diffusion Composition
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee3c592-d739-4461-ae0c-81d25979e657 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision mamba: A multi-grained state space model with self-evolution regularization for offline rl
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9161b203-157a-43c2-a75a-e6c58b3bf523 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning latent plans from play
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7760b84f-ef4a-4c07-b158-dd9ed90e7346 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization How Far I'll Go: Offline Goal-Conditioned Reinforcement Learning via $f$-Advantage Regression
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c121fe31-7789-4cdd-bd85-6da077e82bd2 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization VIP : Towards universal visual reward and representation via value-implicit pre-training
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81535569-6821-4cf1-b589-a68adfbf95cf · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Human-level control through deep reinforcement learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b836f8df-5d17-418a-9393-f6a1a8ccaa13 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d110249c-705f-4d70-a3d4-ceec22004d39 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Asymmetric least squares estimation and testing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6fff31ef-0216-4c08-a905-3cc793992213 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49179c0b-5541-4a55-b2de-51de3d7635b3 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Hiql: Offline goal-conditioned rl with latent states as actions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5f5575a-0a9d-432b-9163-2d9147e6f3df · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Foundation Policies with Hilbert Representations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6d3909-af36-4f36-b8d2-07deef469bd8 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Ogbench: Benchmarking offline goal-conditioned rl
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 03cde593-9e7a-4b38-b8c0-d2b09ffe2fe9 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization A survey on offline reinforcement learning: Taxonomy, review, and open problems
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 738e246b-af48-488e-bd3d-3f9c6928ce3d · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-Conditioned Imitation Learning using Score-based Diffusion Policies
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0047489f-1d2b-4ace-91ba-3d824f4a0da4 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stochastic backpropagation and approximate inference in deep generative models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f43cd76c-e99b-47ef-a3bc-6bd55c4b6680 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Outcome-driven reinforcement learning via variational inference
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d75398f-09e4-484f-a1a0-d8b10e8421f6 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinforcement learning upside down: Don't predict rewards -- just map them to actions, 2020
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0ad97e9-02cb-4edd-8211-cb21fa4dab4d · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rapid exploration for open-world navigation with latent goal models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfc8bf9e-530b-400e-a12c-6bd0d5fc4e6f · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Score models for offline goal-conditioned reinforcement learning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9e70b2b-72cf-44f6-996f-cfd0632dad84 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Geoadditive expectile regression
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9c00165-7d4f-4aeb-9178-49005ea0b73e · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Learning structured output representation using deep conditional generative models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cf407e-cc0a-4685-ae51-73d4d982f234 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Gymnasium (mar 2023), 2023
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15c15ca4-5ecf-46a0-abb1-078650721ce1 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Deep Reinforcement Learning and the Deadly Triad
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e47b5a1-f2c6-4aec-b59c-e20c712aadcc · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Attention is all you need
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20b3ea7-d6cd-4f80-a058-c02026c8f25b · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization GOP lan: Goal-conditioned offline reinforcement learning by planning with learned models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34fa6ff0-f0f9-4b0f-bac5-f5c5f9efce2b · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Optimal goal-reaching reinforcement learning via quasimetric learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4064caf9-0460-452a-a414-d3ea921924ee · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Critic-guided decision transformer for offline reinforcement learning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f240bd1-3f82-44a7-900f-bb0dcb6eb6ec · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Supported policy optimization for offline reinforcement learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e8e42e3d-1cdb-42c8-9228-c14463d124a2 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Elastic Decision Transformer
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8c62b3-2b08-4809-9704-568102ecae9d · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 158ff2de-8365-47d0-809c-dfad1cc405e2 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0fa5c6-5581-4b25-87fb-013fcef04677 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization What is essential for unseen goal generalization of offline goal-conditioned rl? In International Conference on Machine Learning, pages 39543--39571
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77db9b91-9fca-46a6-9296-56fca0933486 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Swapped goal-conditioned offline reinforcement learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8dd9a5-e9c1-4d76-af4a-b181000ab2fb · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Breadth-first exploration on adaptive grid for reinforcement learning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbcd7cd-bb20-4774-be4e-c2841fc8cbee · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Goal-conditioned predictive coding for offline reinforcement learning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b155832e-295f-432b-bfcd-1de1323f54b8 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b355547-cf41-4fd4-860b-354833c8961b · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Contrastive difference predictive coding
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1007d93e-a759-4670-a9ae-ce0f34545b99 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Online decision transformer
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2de92e6a-ec70-4159-9a24-cf51af5f8e47 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Behavior Proximal Policy Optimization
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1267e842-1074-4441-b705-b7277e3483cd · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Reinformer: Max-Return Sequence Modeling for Offline RL
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad6054d5-8274-4dd4-b327-6f158fb0db69 · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Revisiting the design choices in max-return sequence modeling, 2025
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc578017-cce3-4fd5-bb09-46cadeecab2d · outbound
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization Maximum entropy inverse reinforcement learning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.