Pith. sign in

Paper Citation Record · LEDGER

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.02034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02034 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:06:50.537588Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 834b2273-932a-4e92-9cc3-494e444867e1 · outbound

This paper cites Machine learning , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.116931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.116931Z digest=sha256:6233d226ff53240d0356deb8239eb78a828e81c4f74135946146adfb280d212b

Observation 12bd851b-be0c-4587-a88e-8e5045eada13 · outbound

This paper cites Machine learning , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.125517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.125517Z digest=sha256:1b288dbd2cf7e97f4eb0765a4958cc18f8fec09b3267f94b7d2aaa83a2e7376d

Observation 91ab2b70-7529-4ade-9d3c-4bb6132cd5af · outbound

This paper cites 1998 , publisher=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 1998 , publisher=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.131429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.131429Z digest=sha256:4c6e7adfe8719f6a7eb8107d792ebcef90fd2a0ae5ebb7996ea4d73a2e6a494b

Observation 8b6cf0aa-46b9-4a19-b996-e39a565683c4 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.139308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.139308Z digest=sha256:0c5ba19f6559c7ae94d243268f0d4577bad0305c93ad1350632a01cd13e7b2d0

Observation debd968b-cac0-4fde-9d45-b8718bb7b4a3 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.640450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.146215Z digest=sha256:67571cbe9dfd34f298828d1b2f722b30ba99ffafc50314feb7210b324fe5c703

Observation b877c83e-c40e-43f4-a674-5abb61ebcb15 · outbound

This paper cites and Singh, Satinder P.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Singh, Satinder P

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.609883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.151699Z digest=sha256:a21187a9389a33663dfa8fe9e9a754e4c04fd44cee76981782108f8b1867562d

Observation 7a8c56f3-2484-4e55-a04c-f82d8f81af14 · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.158210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.158210Z digest=sha256:7a42bdcb022801fc3e2d86b12ccc8996080f6491e3446c0b7d0be5ec9520b979

Observation e8636fba-5dfa-4f57-aa4a-d8842223c53d · outbound

This paper cites International conference on machine learning , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International conference on machine learning , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.163562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.163562Z digest=sha256:214afd565037f3622fda54ddbcd2e93c4d5b1f2c3db8254877fda50f6322bb02

Observation 86945018-3bd9-4ed1-9d3b-8a656a86de69 · outbound

This paper cites Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.170813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.170813Z digest=sha256:c9eef8795125815e2ddc0edefc7cd8f06607f6d578535658bab9ffda7806e331

Observation e25c9158-9fed-4f50-b4a8-a1f00e1ef290 · outbound

This paper cites Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:06:50.980075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.178184Z digest=sha256:4b25f1ffa8bcd3d292971e349d09ae3badda447030c763406e12c070e3a5d38d

Observation 66e53fed-4356-44a0-808e-a68d3b469e54 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.184883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.184883Z digest=sha256:a5ec24ae1f6094688ca20244488e5e97b93a9353a34275146d74f45c5ef5dc96

Observation 20012c6e-ff60-40d2-8d45-2a8968370b20 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.191923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.191923Z digest=sha256:e74f2b64042e84f9f21dac4b244990650229209ec5b72ad355d9fec4f457e4ae

Observation 419d81fd-c36a-4ca0-8778-2907a4590fd6 · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Econometrica: Journal of the Econometric Society , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.199804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.199804Z digest=sha256:02902402a683a8539a81b00cef5b90c3718be545f001ae6c3363c660f9ebb35f

Observation 68355baa-5fdd-4f87-a493-481468fe76aa · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.206054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.206054Z digest=sha256:081c97fab6d49fd2ef9fd0b2eeed2df3681c2810cff0fea2d9ea9c66d60fca10

Observation cb32992d-f383-4c1f-841e-67b1e79c1627 · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Forty-second International Conference on Machine Learning , year=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.492116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.211907Z digest=sha256:9ed9b2cb27807474c2105def11145b8fb429bd54f964137507867ac590b56e06

Observation dab51c5b-b274-4c6b-a9a5-8c193c8191d0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.226614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.226614Z digest=sha256:6871dd9fcc18d0a179518020b078c29081b3e3fdb7c02f9f989f1de5fda60a6f

Observation 496826eb-b602-4b24-8e85-a4cf9ca15887 · outbound

This paper cites International Conference on Learning Representations , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.232980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.232980Z digest=sha256:64fefa73f7869a9d446d20616b983f4fb56d963185d24a11f179829adf1642e6

Observation a9994f8e-f09f-4a61-92c5-328bb4e16d45 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.239968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.239968Z digest=sha256:0d32cbb720bf6a3e2d5cdf79dcc97d2a5ad919c81de1dd30f32e8caf6b4b5b9b

Observation 74d2f582-6ada-4aa0-a076-048f86a3b609 · outbound

This paper cites Revisiting the Minimalist Approach to Offline Reinforcement Learning , url =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting the Minimalist Approach to Offline Reinforcement Learning , url =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.394137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.245634Z digest=sha256:efb1a0b727446d9c5120f4ecb5e9d7b9924c5ba14ea2652546c87775dc321ada

Observation 1a4c69b9-d45d-4b90-b579-d46226570b7c · outbound

This paper cites Flow Matching for Generative Modeling.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Flow Matching for Generative Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.252484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.252484Z digest=sha256:b699ea34ebe59298f9739fe8e8455c2ab8df261a47227c4c69a5da53df9b272c

Observation da4a590a-0071-4cb1-856d-543180428c75 · outbound

This paper cites The International Journal of Robotics Research , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The International Journal of Robotics Research , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.259885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.259885Z digest=sha256:e036636b8e095bd7ef149ce347735722651eb2e6598ffc7817b8defd079f2880

Observation d44ecb98-a287-4476-b206-a7317ecaea3c · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirteenth International Conference on Learning Representations , year=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.343344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.272737Z digest=sha256:0e8f225458392abff244f5f945049047cdec78b0dc358f54634f689fb4fb0966

Observation 09f93716-0152-4043-acce-878b7557a259 · outbound

This paper cites 2026 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.307860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.278592Z digest=sha256:e3d76c4553dc336b2990ce34c5f19f0113e1f2518abf650c25c4f0e7214b77d5

Observation f64bd37b-89ba-43af-a3ed-b828eaf18312 · outbound

This paper cites Value-Based Deep.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Value-Based Deep

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.269376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.286476Z digest=sha256:9e3f32a932036528ddd3d036dd101f749d4ceef3cd74058f3f4e9c847329f88e

Observation c8cacb51-39ac-41e3-a043-5315563beba8 · outbound

This paper cites The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.238878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.291836Z digest=sha256:f106839aad49142fe1879f28469890836808e41c974c8aa7269ad965bf74a374

Observation 568ff334-8390-46a4-aad1-9113bf014708 · outbound

This paper cites 2025 , url=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2025 , url=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.299033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.299033Z digest=sha256:2a9aeed20a399e6508cdc851eb8a39e46b0dae6d3a40d1702002c6df7bb174e1

Observation c18b6e6e-fa65-438d-90cc-923ebd1e319d · outbound

This paper cites an unresolved cited work.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-15T15:06:50.607406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.304764Z digest=sha256:0f2b17b1aea8fc96b45432532437d0e6ac26940dff8a65066107e4bbc058bb0c

Observation d987d442-2f46-478f-a807-7d643c09c539 · outbound

This paper cites , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning , title =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.193012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.310504Z digest=sha256:f9a2f55164e5afe9fb3e26734bbc9dacdf37571988dbe5a8bad155fd8f0c246e

Observation d2c3ccf7-3683-4645-9b9c-429042cf6316 · outbound

This paper cites 2024 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2024 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.166940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.316061Z digest=sha256:bf6cbaea724f36ac1d56e8c5715dadb09f9f618ed36a2a10f8a417b5fba806b3

Observation ccb1c3ed-fea9-4510-8a41-81a05f66daf8 · outbound

This paper cites 2026 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.141488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.321895Z digest=sha256:c48c57de440e0233b90a1aaa8f0244e6d896add57bf8ccf9da62fe4d091232fd

Observation 4fb82770-ea45-40ee-bee7-975f24d000fb · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.327384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.327384Z digest=sha256:4397cb724e4844795b62443b8976512c7d276164761a8c5343dd80fafc4de1eb

Observation 90f1bccf-156c-42cf-b501-0f374521c707 · outbound

This paper cites Proceedings of the 1993 connectionist models summer school , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 1993 connectionist models summer school , pages=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.106774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.332550Z digest=sha256:20a1fd89c825f0d174e5ade53c033c72321bd721e60f7bb9f6a3f6c97f45dd61

Observation 0b1eadec-d416-4fd4-adab-c085986120e4 · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.062103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.337580Z digest=sha256:778be6287661bb5ca0702ab9a2d32cff4a12429a0f0f48c44933c918624b88ff

Observation 4ec39e56-1621-4dd7-a174-a7a0c921eaac · outbound

This paper cites Sur les op.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Sur les op

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.342760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.342760Z digest=sha256:e0ac444ac56f55f8ff3f1c27cc6ec56bdfa40571f94c6551dd4af5b6078c49e7

Observation a5460b76-d901-4b88-89f4-4135348ed60a · outbound

This paper cites 2014 , publisher=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2014 , publisher=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.347855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.347855Z digest=sha256:887723018ae8f7e29f2112c9d2e9c3c8e2202e039ffb4468478eb0e89ad11f3a

Observation b0b8dddb-085d-4aa2-8ea1-4c3a6c988b37 · outbound

This paper cites Generalized quantiles as risk measures , journal =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Generalized quantiles as risk measures , journal =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.352937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.352937Z digest=sha256:4f3cb644b37b507e37e3c2a26f236ff3fa4dab12a3496aa8cb079e0bed8e8054

Observation d657fc87-5285-4c5d-935c-3156b5ec8352 · outbound

This paper cites 2013 , month =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2013 , month =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.357612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.357612Z digest=sha256:5cdc73157a733f75774bd413149423a2130b5b28ffad3140abb2ebd88f3080ef

Observation 5e126664-4372-44f1-8a2f-6d1917e4f7c5 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.362872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.362872Z digest=sha256:b145ec45bfaa19a5116773f4b4f973f1b664c11b62b4e85866c50cf4507deba0

Observation 51363595-7795-464a-9960-2014b44e4987 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 36th International Conference on Machine Learning (ICML) , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.982027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.368514Z digest=sha256:095c9aaaa560e39d001c34718aadbbde8977aecc8ea55d7af7fa53d7e59d9278

Observation 04dce202-6b9d-4b43-81c4-a6c53edf6386 · outbound

This paper cites Advances in Neural Information Processing Systems 32 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 32 (NeurIPS) , year =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.952488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.373398Z digest=sha256:b46136fb94b4e494c0484f095aff78016c40af5bbb0ba41361803555816a0b43

Observation 584a5242-24fa-4346-9059-c24d2a2f407d · outbound

This paper cites Advances in Neural Information Processing Systems 33 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 33 (NeurIPS) , year =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.923532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.378202Z digest=sha256:0f6112b9d9380963bd7057effccca6ce86004e5a0a0cdd6abdc6a3e82f84f773

Observation e44361ea-def0-4f0f-af68-0cbfb2be80b8 · outbound

This paper cites Advances in Neural Information Processing Systems 34 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.898227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.384263Z digest=sha256:c832649503490e4b4db03162681ec0393ad4fbfa34c0e219d9a20cda23085295

Observation 7febf1af-9dde-4909-a21f-586a98d3f19a · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.875190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.389719Z digest=sha256:2bddac47bb4add8dd7d872c22a806a2c8c0ba595b79c4402e69c117290584328

Observation f191590c-e1dc-4e78-bf6c-f7a0d6f0489f · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.833757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.395788Z digest=sha256:809bb7207ecb8894187788fa0e68efa7384f1464363744dcb09f2db414ada4d3

Observation 89319c69-97f7-4606-a6ce-d3faec6eef93 · outbound

This paper cites Advances in Neural Information Processing Systems 35 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 35 (NeurIPS) , year =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.807942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.400726Z digest=sha256:f36c479a09d501b0babec2badb3341ffdb9dd30862117e012f0a75b0091937a0

Observation ac5e6bf4-0354-4fbc-b3dd-dd336d88f302 · outbound

This paper cites Offline Reinforcement Learning as Anti-Exploration , booktitle =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning as Anti-Exploration , booktitle =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.773722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.406133Z digest=sha256:765e7a21b21dfa196b741e67f46455e794b5316e896b71e3a5da750ea03ac68b

Observation b2cdb4a7-f0f9-472e-a7a4-112f3396bbc7 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 40th International Conference on Machine Learning (ICML) , year =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.746685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.411207Z digest=sha256:aa736f1748f27ce6bbfcde5b2f1b071c57a6b9146ce795390d8a04db3d64875f

Observation ab042adc-a7c6-412f-b1f8-2eb7a29ffcfe · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 38th International Conference on Machine Learning (ICML) , year =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.715621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.416116Z digest=sha256:94f3f3072ad5c23ccb3249245d0c0ab4b7dd72abd1034fc9ac65ffa9f5cabc19

Observation 5842a013-6e7a-4a2f-bc18-2744fa0eb4f2 · outbound

This paper cites Advances in Neural Information Processing Systems 34 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.680986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.422817Z digest=sha256:7477ffd822c48067e865c75be3be57ef44d414080179f66af8287a32511cfc35

Observation 3a063cd4-19df-4bcd-8aa6-9244456288f8 · outbound

This paper cites Proceedings of the 35th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 35th International Conference on Machine Learning (ICML) , year =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.648641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.427509Z digest=sha256:3b174a79f49d0fded64c8b4d9f5b4451b3527f857d607e7ad0f7044d1ead288e

Observation f72edb3c-a79a-46cc-80a0-7570cc6b2ef4 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 37th International Conference on Machine Learning (ICML) , year =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.616154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.432448Z digest=sha256:171bf188ffffaa2b7888468a5d8bced7c8991079e4ef8198a40a40e967060b2f

Observation 4cfd3ffa-40db-4761-9787-4ac0b52919a1 · outbound

This paper cites Revisiting.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.587948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.437973Z digest=sha256:0b5d309f20b4919797dbc44b16b7604709524b5b7048fa31bc3e0c846016f29a

Observation d4f97450-5921-4a03-8ad3-c1241a9a3717 · outbound

This paper cites Advances in Neural Information Processing Systems 38 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 38 (NeurIPS) , year =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.552686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.442654Z digest=sha256:76fce899cd66df500ccf30f543b81a21392d129e44aa5c0bc4c0c97aa409fc29

Observation c4301007-bce7-4822-b88c-31f0c0d25283 · outbound

This paper cites and Munos, R.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Munos, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.513390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.447393Z digest=sha256:f96f37d1cdeddbbf08f71be80b812684deee0160f5b46ac5121d8e9cdff940cf

Observation e47060d1-6c2b-4120-86cf-347cfc7966fd · outbound

This paper cites Statistics and Samples in Distributional Reinforcement Learning , booktitle =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Statistics and Samples in Distributional Reinforcement Learning , booktitle =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.456611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.452227Z digest=sha256:35fb151c5ad63a68ad7503d8ed4c28c7d54f0f8c0676fab438a5aaa587b82feb

Observation e848eb86-a4b3-430a-aa00-676d2b8aa517 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.430630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.456880Z digest=sha256:8a713dcbe0f53cf482b7829937dacf791c8921d3af9caf67b85718a828b62961

Observation 33ded815-72d9-416e-a63f-8e8804e547fc · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.403649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.466583Z digest=sha256:04d45d6e503257a49024292a475c006dea4d1343c30b070ef48d9d83f800b6b0

Observation 670b1a81-6766-4050-ba76-1cf70d38fee3 · outbound

This paper cites and Zhou, Mingyuan , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Zhou, Mingyuan , title =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.382181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.473296Z digest=sha256:19019ba5e45a6630fe4168d3160c923c51aae5a7e39b9e97b1a7e096335b8540

Observation a72d2690-b344-4752-ae67-70ee42ca6f48 · outbound

This paper cites and Kumar, Vikash and Levine, Sergey and Finn, Chelsea , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Kumar, Vikash and Levine, Sergey and Finn, Chelsea , title =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.358454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.478163Z digest=sha256:7b9ef1de865a49432bba543ce98755c222e2e293467cdd4ad6c05b0d88323b87

Observation a67bb99a-ebe4-4c44-8289-3586cbede48f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.483090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.483090Z digest=sha256:bb05474df8556ad795343c013079c03125bad5c1d9d5ac65acbf6d3221bcc601

Observation b159dcac-838a-448f-9c8b-ea97161450e4 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 36 (NeurIPS) , year =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.325830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.493387Z digest=sha256:1459198d0cadf6758a8bb0d12eb99a7525f46c448589ecd39e9647d0bbab2dba

Observation 9a7f8a6a-1557-4061-8488-4eaf0592c0d5 · outbound

This paper cites and Smith, Laura and Kostrikov, Ilya and Levine, Sergey , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Smith, Laura and Kostrikov, Ilya and Levine, Sergey , title =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.276547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.499712Z digest=sha256:a6f5b75aff456fafb0e2be93270d5728ecfecaabd073887d86539479db8efa3f

Observation 174d633a-e74d-4d99-9707-6fb712d0387c · outbound

This paper cites and Bellemare, Marc G.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Bellemare, Marc G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.247226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.505717Z digest=sha256:5e0ddd15c1070dc2f6dde089c60d71b9b2e7b63968cca33815da0735ead545e7

Observation 6b04cae4-6669-4745-a134-b74f16b7b5cc · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.207655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.511800Z digest=sha256:91e84ab24c8508eb6c1e7dd404d340fd65afbdc745cd7f037face4cde1d5566c

Observation 4e9b1ee6-c62b-4dea-acdb-6b592f06a8f2 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.156893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.517833Z digest=sha256:563a52876a5ae2c207eebae1ee42b710095949c5ad171f0fb7e7868ae6cb758e

Observation ab82142d-8529-443d-84e3-95f35dbfd25b · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.120731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.523422Z digest=sha256:ffd51884c059fc27622a194af04f4761bb57b8e29b2d084dca8c7dbcb698b5e0

Observation 8e1bd23b-ad48-41e2-a5c0-b5d1b711de30 · outbound

This paper cites Advances in Neural Information Processing Systems 37 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 37 (NeurIPS) , year =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.090106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.530052Z digest=sha256:059666b5af2ce68647ddb565f1e1332cdd09a29716c35da70e0b498ee9921b2d

Observation f33cebe7-29a0-40cf-ac69-445cb37dc799 · outbound

This paper cites Atti del Congresso Internazionale dei Matematici, Bologna , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Atti del Congresso Internazionale dei Matematici, Bologna , volume =

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.060220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.537588Z digest=sha256:3df06f967ce06159d54add38897f74f518d9a32003e7809375aa8c016c197de7

Pith citing papers

No inbound Pith citation observations are available.