Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:06:50.537588Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.02034.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:06:50.537588Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 834b2273-932a-4e92-9cc3-494e444867e1 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bd851b-be0c-4587-a88e-8e5045eada13 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ab2b70-7529-4ade-9d3c-4bb6132cd5af · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 1998 , publisher=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6cf0aa-46b9-4a19-b996-e39a565683c4 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation debd968b-cac0-4fde-9d45-b8718bb7b4a3 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b877c83e-c40e-43f4-a674-5abb61ebcb15 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Singh, Satinder P
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7a8c56f3-2484-4e55-a04c-f82d8f81af14 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8636fba-5dfa-4f57-aa4a-d8842223c53d · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International conference on machine learning , pages=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86945018-3bd9-4ed1-9d3b-8a656a86de69 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25c9158-9fed-4f50-b4a8-a1f00e1ef290 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 66e53fed-4356-44a0-808e-a68d3b469e54 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20012c6e-ff60-40d2-8d45-2a8968370b20 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419d81fd-c36a-4ca0-8778-2907a4590fd6 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Econometrica: Journal of the Econometric Society , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68355baa-5fdd-4f87-a493-481468fe76aa · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb32992d-f383-4c1f-841e-67b1e79c1627 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Forty-second International Conference on Machine Learning , year=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dab51c5b-b274-4c6b-a9a5-8c193c8191d0 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496826eb-b602-4b24-8e85-a4cf9ca15887 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations , volume=
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9994f8e-f09f-4a61-92c5-328bb4e16d45 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d2f582-6ada-4aa0-a076-048f86a3b609 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting the Minimalist Approach to Offline Reinforcement Learning , url =
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a4c69b9-d45d-4b90-b579-d46226570b7c · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Flow Matching for Generative Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4a590a-0071-4cb1-856d-543180428c75 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The International Journal of Robotics Research , volume=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44ecb98-a287-4476-b206-a7317ecaea3c · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirteenth International Conference on Learning Representations , year=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 09f93716-0152-4043-acce-878b7557a259 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f64bd37b-89ba-43af-a3ed-b828eaf18312 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Value-Based Deep
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c8cacb51-39ac-41e3-a043-5315563beba8 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 568ff334-8390-46a4-aad1-9113bf014708 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2025 , url=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18b6e6e-fa65-438d-90cc-923ebd1e319d · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d987d442-2f46-478f-a807-7d643c09c539 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning , title =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d2c3ccf7-3683-4645-9b9c-429042cf6316 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2024 , eprint=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ccb1c3ed-fea9-4510-8a41-81a05f66daf8 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fb82770-ea45-40ee-bee7-975f24d000fb · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f1bccf-156c-42cf-b501-0f374521c707 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 1993 connectionist models summer school , pages=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0b1eadec-d416-4fd4-adab-c085986120e4 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4ec39e56-1621-4dd7-a174-a7a0c921eaac · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Sur les op
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5460b76-d901-4b88-89f4-4135348ed60a · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2014 , publisher=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b8dddb-085d-4aa2-8ea1-4c3a6c988b37 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Generalized quantiles as risk measures , journal =
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d657fc87-5285-4c5d-935c-3156b5ec8352 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2013 , month =
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e126664-4372-44f1-8a2f-6d1917e4f7c5 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51363595-7795-464a-9960-2014b44e4987 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 36th International Conference on Machine Learning (ICML) , year =
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 04dce202-6b9d-4b43-81c4-a6c53edf6386 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 32 (NeurIPS) , year =
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 584a5242-24fa-4346-9059-c24d2a2f407d · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 33 (NeurIPS) , year =
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e44361ea-def0-4f0f-af68-0cbfb2be80b8 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7febf1af-9dde-4909-a21f-586a98d3f19a · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f191590c-e1dc-4e78-bf6c-f7a0d6f0489f · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 89319c69-97f7-4606-a6ce-d3faec6eef93 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 35 (NeurIPS) , year =
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ac5e6bf4-0354-4fbc-b3dd-dd336d88f302 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning as Anti-Exploration , booktitle =
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b2cdb4a7-f0f9-472e-a7a4-112f3396bbc7 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 40th International Conference on Machine Learning (ICML) , year =
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab042adc-a7c6-412f-b1f8-2eb7a29ffcfe · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 38th International Conference on Machine Learning (ICML) , year =
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5842a013-6e7a-4a2f-bc18-2744fa0eb4f2 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a063cd4-19df-4bcd-8aa6-9244456288f8 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 35th International Conference on Machine Learning (ICML) , year =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f72edb3c-a79a-46cc-80a0-7570cc6b2ef4 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 37th International Conference on Machine Learning (ICML) , year =
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4cfd3ffa-40db-4761-9787-4ac0b52919a1 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d4f97450-5921-4a03-8ad3-c1241a9a3717 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 38 (NeurIPS) , year =
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c4301007-bce7-4822-b88c-31f0c0d25283 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Munos, R
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e47060d1-6c2b-4120-86cf-347cfc7966fd · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Statistics and Samples in Distributional Reinforcement Learning , booktitle =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e848eb86-a4b3-430a-aa00-676d2b8aa517 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33ded815-72d9-416e-a63f-8e8804e547fc · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 670b1a81-6766-4050-ba76-1cf70d38fee3 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Zhou, Mingyuan , title =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a72d2690-b344-4752-ae67-70ee42ca6f48 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Kumar, Vikash and Levine, Sergey and Finn, Chelsea , title =
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a67bb99a-ebe4-4c44-8289-3586cbede48f · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b159dcac-838a-448f-9c8b-ea97161450e4 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 36 (NeurIPS) , year =
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9a7f8a6a-1557-4061-8488-4eaf0592c0d5 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Smith, Laura and Kostrikov, Ilya and Levine, Sergey , title =
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 174d633a-e74d-4d99-9707-6fb712d0387c · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Bellemare, Marc G
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6b04cae4-6669-4745-a134-b74f16b7b5cc · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e9b1ee6-c62b-4dea-acdb-6b592f06a8f2 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab82142d-8529-443d-84e3-95f35dbfd25b · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8e1bd23b-ad48-41e2-a5c0-b5d1b711de30 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 37 (NeurIPS) , year =
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f33cebe7-29a0-40cf-ac69-445cb37dc799 · outbound
Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Atti del Congresso Internazionale dei Matematici, Bologna , volume =
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.