Pith. sign in

Paper Citation Record · LEDGER

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2608.02034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02034 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:06:50.537588Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 834b2273-932a-4e92-9cc3-494e444867e1 · outbound

This paper cites Machine learning , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.116931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.116931Z digest=sha256:6233d226ff53240d0356deb8239eb78a828e81c4f74135946146adfb280d212b

Observation 12bd851b-be0c-4587-a88e-8e5045eada13 · outbound

This paper cites Machine learning , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Machine learning , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.125517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.125517Z digest=sha256:1b288dbd2cf7e97f4eb0765a4958cc18f8fec09b3267f94b7d2aaa83a2e7376d

Observation 91ab2b70-7529-4ade-9d3c-4bb6132cd5af · outbound

This paper cites 1998 , publisher=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 1998 , publisher=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.131429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.131429Z digest=sha256:4c6e7adfe8719f6a7eb8107d792ebcef90fd2a0ae5ebb7996ea4d73a2e6a494b

Observation 8b6cf0aa-46b9-4a19-b996-e39a565683c4 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.139308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.139308Z digest=sha256:0c5ba19f6559c7ae94d243268f0d4577bad0305c93ad1350632a01cd13e7b2d0

Observation debd968b-cac0-4fde-9d45-b8718bb7b4a3 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.640450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.146215Z digest=sha256:4665dac9c42c104a1237db1b2e355723c9322294eee98addd2b6cdb0b5a471da

Observation b877c83e-c40e-43f4-a674-5abb61ebcb15 · outbound

This paper cites and Singh, Satinder P.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Singh, Satinder P

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.609883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.151699Z digest=sha256:686a12b476198a176441c01832255480965f7070275e9ca37d6eeb9d727ea659

Observation 7a8c56f3-2484-4e55-a04c-f82d8f81af14 · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.158210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.158210Z digest=sha256:7a42bdcb022801fc3e2d86b12ccc8996080f6491e3446c0b7d0be5ec9520b979

Observation e8636fba-5dfa-4f57-aa4a-d8842223c53d · outbound

This paper cites International conference on machine learning , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International conference on machine learning , pages=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.163562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.163562Z digest=sha256:214afd565037f3622fda54ddbcd2e93c4d5b1f2c3db8254877fda50f6322bb02

Observation 86945018-3bd9-4ed1-9d3b-8a656a86de69 · outbound

This paper cites Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Learning to Play in a Day: Faster Deep Reinforcement Learning by Optimality Tightening

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.170813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.170813Z digest=sha256:c9eef8795125815e2ddc0edefc7cd8f06607f6d578535658bab9ffda7806e331

Observation e25c9158-9fed-4f50-b4a8-a1f00e1ef290 · outbound

This paper cites Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:06:50.980075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.178184Z digest=sha256:d28d1fc135bc987833454afed12691d704aee7c879ee028464d57a167e7da34b

Observation 66e53fed-4356-44a0-808e-a68d3b469e54 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.184883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.184883Z digest=sha256:a5ec24ae1f6094688ca20244488e5e97b93a9353a34275146d74f45c5ef5dc96

Observation 20012c6e-ff60-40d2-8d45-2a8968370b20 · outbound

This paper cites Extreme Q-Learning: MaxEnt RL without Entropy.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Extreme Q-Learning: MaxEnt RL without Entropy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.191923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.191923Z digest=sha256:e74f2b64042e84f9f21dac4b244990650229209ec5b72ad355d9fec4f457e4ae

Observation 419d81fd-c36a-4ca0-8778-2907a4590fd6 · outbound

This paper cites Econometrica: Journal of the Econometric Society , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Econometrica: Journal of the Econometric Society , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.199804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.199804Z digest=sha256:02902402a683a8539a81b00cef5b90c3718be545f001ae6c3363c660f9ebb35f

Observation 68355baa-5fdd-4f87-a493-481468fe76aa · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.206054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.206054Z digest=sha256:081c97fab6d49fd2ef9fd0b2eeed2df3681c2810cff0fea2d9ea9c66d60fca10

Observation cb32992d-f383-4c1f-841e-67b1e79c1627 · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Forty-second International Conference on Machine Learning , year=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.492116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.211907Z digest=sha256:143a4dbd3154ac574a830eec493153e0d20d2e056921ed8d5d7a4396212037ed

Observation dab51c5b-b274-4c6b-a9a5-8c193c8191d0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.226614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.226614Z digest=sha256:6871dd9fcc18d0a179518020b078c29081b3e3fdb7c02f9f989f1de5fda60a6f

Observation 496826eb-b602-4b24-8e85-a4cf9ca15887 · outbound

This paper cites International Conference on Learning Representations , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.232980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.232980Z digest=sha256:64fefa73f7869a9d446d20616b983f4fb56d963185d24a11f179829adf1642e6

Observation a9994f8e-f09f-4a61-92c5-328bb4e16d45 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.239968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.239968Z digest=sha256:0d32cbb720bf6a3e2d5cdf79dcc97d2a5ad919c81de1dd30f32e8caf6b4b5b9b

Observation 74d2f582-6ada-4aa0-a076-048f86a3b609 · outbound

This paper cites Revisiting the Minimalist Approach to Offline Reinforcement Learning , url =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting the Minimalist Approach to Offline Reinforcement Learning , url =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.394137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.245634Z digest=sha256:2dfa481a604ea224664a9fd3fc3a618d7e9e308da5a8dd09d82f77ce73baa382

Observation 1a4c69b9-d45d-4b90-b579-d46226570b7c · outbound

This paper cites Flow Matching for Generative Modeling.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Flow Matching for Generative Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.252484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.252484Z digest=sha256:b699ea34ebe59298f9739fe8e8455c2ab8df261a47227c4c69a5da53df9b272c

Observation da4a590a-0071-4cb1-856d-543180428c75 · outbound

This paper cites The International Journal of Robotics Research , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The International Journal of Robotics Research , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.259885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.259885Z digest=sha256:e036636b8e095bd7ef149ce347735722651eb2e6598ffc7817b8defd079f2880

Observation d44ecb98-a287-4476-b206-a7317ecaea3c · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirteenth International Conference on Learning Representations , year=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.343344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.272737Z digest=sha256:41c66de2f36d6a33b885b626f6bcb717d3bcfc1356558ed579ce420069d305be

Observation 09f93716-0152-4043-acce-878b7557a259 · outbound

This paper cites 2026 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.307860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.278592Z digest=sha256:6c98f86c78e865fcf91da31708cf6b700e7cda5b83491f548f89270b82861854

Observation f64bd37b-89ba-43af-a3ed-b828eaf18312 · outbound

This paper cites Value-Based Deep.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Value-Based Deep

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.269376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.286476Z digest=sha256:1b645e7c34b949c12f8e127ca2fb850abb4bd72f125c9659678410eaa9109c07

Observation c8cacb51-39ac-41e3-a043-5315563beba8 · outbound

This paper cites The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.238878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.291836Z digest=sha256:120494f5e803d1b389cf60d5c1d6ad08350440c2023e73c4503b2aeebc67499c

Observation 568ff334-8390-46a4-aad1-9113bf014708 · outbound

This paper cites 2025 , url=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2025 , url=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.299033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.299033Z digest=sha256:2a9aeed20a399e6508cdc851eb8a39e46b0dae6d3a40d1702002c6df7bb174e1

Observation c18b6e6e-fa65-438d-90cc-923ebd1e319d · outbound

This paper cites an unresolved cited work.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-15T15:06:50.607406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.304764Z digest=sha256:3c1273ab205fcb34272b82f5969048b61cd700b78a2196740c763a8e655db2f4

Observation d987d442-2f46-478f-a807-7d643c09c539 · outbound

This paper cites , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning , title =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.193012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.310504Z digest=sha256:9430e3b0f6f2059f85950ef701a777eab9cbec6606ec88117d35cd7eea625738

Observation d2c3ccf7-3683-4645-9b9c-429042cf6316 · outbound

This paper cites 2024 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2024 , eprint=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.166940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.316061Z digest=sha256:b5174b2dde7d3ee7576097c3af3a280dffb771051b2b61ca74c1ffd6ce0fc7b7

Observation ccb1c3ed-fea9-4510-8a41-81a05f66daf8 · outbound

This paper cites 2026 , eprint=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2026 , eprint=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.141488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.321895Z digest=sha256:875f688768c173fca09d828675144f19088c49ab20f076e9ac8bc605a850ec49

Observation 4fb82770-ea45-40ee-bee7-975f24d000fb · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.327384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.327384Z digest=sha256:4397cb724e4844795b62443b8976512c7d276164761a8c5343dd80fafc4de1eb

Observation 90f1bccf-156c-42cf-b501-0f374521c707 · outbound

This paper cites Proceedings of the 1993 connectionist models summer school , pages=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 1993 connectionist models summer school , pages=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.106774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.332550Z digest=sha256:7715751018ae4a3f20523b150fcd894980deeb57be77a6645eef6adf4942bdd0

Observation 0b1eadec-d416-4fd4-adab-c085986120e4 · outbound

This paper cites Advances in neural information processing systems , volume=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in neural information processing systems , volume=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:52.062103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.337580Z digest=sha256:1f19e5735224a2032a9d7ff3105232fe17d410fcd93f385dd1dc2b0b020447b5

Observation 4ec39e56-1621-4dd7-a174-a7a0c921eaac · outbound

This paper cites Sur les op.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Sur les op

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.342760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.342760Z digest=sha256:e0ac444ac56f55f8ff3f1c27cc6ec56bdfa40571f94c6551dd4af5b6078c49e7

Observation a5460b76-d901-4b88-89f4-4135348ed60a · outbound

This paper cites 2014 , publisher=.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2014 , publisher=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.347855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.347855Z digest=sha256:887723018ae8f7e29f2112c9d2e9c3c8e2202e039ffb4468478eb0e89ad11f3a

Observation b0b8dddb-085d-4aa2-8ea1-4c3a6c988b37 · outbound

This paper cites Generalized quantiles as risk measures , journal =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Generalized quantiles as risk measures , journal =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.352937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.352937Z digest=sha256:4f3cb644b37b507e37e3c2a26f236ff3fa4dab12a3496aa8cb079e0bed8e8054

Observation d657fc87-5285-4c5d-935c-3156b5ec8352 · outbound

This paper cites 2013 , month =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 2013 , month =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.357612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.357612Z digest=sha256:5cdc73157a733f75774bd413149423a2130b5b28ffad3140abb2ebd88f3080ef

Observation 5e126664-4372-44f1-8a2f-6d1917e4f7c5 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.362872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.362872Z digest=sha256:b145ec45bfaa19a5116773f4b4f973f1b664c11b62b4e85866c50cf4507deba0

Observation 51363595-7795-464a-9960-2014b44e4987 · outbound

This paper cites Proceedings of the 36th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 36th International Conference on Machine Learning (ICML) , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.982027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.368514Z digest=sha256:c4b8e0aa44e9545af771e136d90524c4b34ca872bdd78d1973b42fce0e059a10

Observation 04dce202-6b9d-4b43-81c4-a6c53edf6386 · outbound

This paper cites Advances in Neural Information Processing Systems 32 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 32 (NeurIPS) , year =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.952488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.373398Z digest=sha256:c547047f994d32b4163a7e8fcc19fb960062330a61c6fa36f9e2c0e943a235d0

Observation 584a5242-24fa-4346-9059-c24d2a2f407d · outbound

This paper cites Advances in Neural Information Processing Systems 33 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 33 (NeurIPS) , year =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.923532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.378202Z digest=sha256:d772fd5909056f54509cb7e0843872e9937c214e2e82da96d63a8916c9dadcb5

Observation e44361ea-def0-4f0f-af68-0cbfb2be80b8 · outbound

This paper cites Advances in Neural Information Processing Systems 34 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.898227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.384263Z digest=sha256:d2659d254986cff96f2b9af88fac9093af28238fba024391385832097f63f027

Observation 7febf1af-9dde-4909-a21f-586a98d3f19a · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.875190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.389719Z digest=sha256:bf3fd5414deda4fdddd1f1e46f53c753dff407e9b471630b5bae9e8197cb4f92

Observation f191590c-e1dc-4e78-bf6c-f7a0d6f0489f · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.833757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.395788Z digest=sha256:58c1edb1b22c2966babf534336901d41f54e33c984404339965e27c03ace2b48

Observation 89319c69-97f7-4606-a6ce-d3faec6eef93 · outbound

This paper cites Advances in Neural Information Processing Systems 35 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 35 (NeurIPS) , year =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.807942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.400726Z digest=sha256:18e1f97ad9d3cfb034336837fc33aa65adb022347580c3c6849b3b7549d9a35a

Observation ac5e6bf4-0354-4fbc-b3dd-dd336d88f302 · outbound

This paper cites Offline Reinforcement Learning as Anti-Exploration , booktitle =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Offline Reinforcement Learning as Anti-Exploration , booktitle =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.773722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.406133Z digest=sha256:87e193f97efe3ae80d3b7be4da7ca2556247e2616dc02bb52265e6d1abdf32df

Observation b2cdb4a7-f0f9-472e-a7a4-112f3396bbc7 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 40th International Conference on Machine Learning (ICML) , year =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.746685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.411207Z digest=sha256:79421d3399f64a9e3d5a9b7b9fe6d7ab735bdef1640050a0f8cc66799f452498

Observation ab042adc-a7c6-412f-b1f8-2eb7a29ffcfe · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 38th International Conference on Machine Learning (ICML) , year =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.715621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.416116Z digest=sha256:4026e923fda6a7e21b7d8e229a7202fb1c41133d144a9cca8ee5e5d3e1c05868

Observation 5842a013-6e7a-4a2f-bc18-2744fa0eb4f2 · outbound

This paper cites Advances in Neural Information Processing Systems 34 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 34 (NeurIPS) , year =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.680986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.422817Z digest=sha256:875ecdb68f007c904e38e57343e016fdf948906257248b4a6bc418c8a958bd7c

Observation 3a063cd4-19df-4bcd-8aa6-9244456288f8 · outbound

This paper cites Proceedings of the 35th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 35th International Conference on Machine Learning (ICML) , year =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.648641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.427509Z digest=sha256:868f9a1a2a38d14e50c28fee2e6866614dfb21cdca32850d5ddf36d5d984d3f9

Observation f72edb3c-a79a-46cc-80a0-7570cc6b2ef4 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning (ICML) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the 37th International Conference on Machine Learning (ICML) , year =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.616154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.432448Z digest=sha256:20ee48c20c280c1c2e6013288aea5d579e1b5cf2ba74bd5467ec5494005d5242

Observation 4cfd3ffa-40db-4761-9787-4ac0b52919a1 · outbound

This paper cites Revisiting.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Revisiting

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.587948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.437973Z digest=sha256:50a65562f6f5ec1f514fc5558a97646747b4ecd0e316fa759e7319a595787bfc

Observation d4f97450-5921-4a03-8ad3-c1241a9a3717 · outbound

This paper cites Advances in Neural Information Processing Systems 38 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 38 (NeurIPS) , year =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.552686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.442654Z digest=sha256:ce0074b09594dd5b716b8a3c8cc1025aaea600e5e9151b87f8b8619a3a999044

Observation c4301007-bce7-4822-b88c-31f0c0d25283 · outbound

This paper cites and Munos, R.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Munos, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.513390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.447393Z digest=sha256:c8242fd072f086c6432de3d8d4841b9ad3ce0691b6675ac111777fcdbc32fb45

Observation e47060d1-6c2b-4120-86cf-347cfc7966fd · outbound

This paper cites Statistics and Samples in Distributional Reinforcement Learning , booktitle =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Statistics and Samples in Distributional Reinforcement Learning , booktitle =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.456611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.452227Z digest=sha256:999d7182600fe7896d41f9f3e6a59fd8fd5465e287a45de6fe84be2b62003efa

Observation e848eb86-a4b3-430a-aa00-676d2b8aa517 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.430630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.456880Z digest=sha256:74e64d6af1a7a7546f0869302bdf119046a90a20a0b065539afde662847596db

Observation 33ded815-72d9-416e-a63f-8e8804e547fc · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.403649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.466583Z digest=sha256:e9646de4eda2ec55d10dc9f8697aada1254e4f395cab8dfdf38dc50a8c65a3c0

Observation 670b1a81-6766-4050-ba76-1cf70d38fee3 · outbound

This paper cites and Zhou, Mingyuan , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Zhou, Mingyuan , title =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.382181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.473296Z digest=sha256:519b4b008ede26dc72ba7c699e8557bb30dcc0cc24a12b54c2ea23dccae72f1f

Observation a72d2690-b344-4752-ae67-70ee42ca6f48 · outbound

This paper cites and Kumar, Vikash and Levine, Sergey and Finn, Chelsea , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Kumar, Vikash and Levine, Sergey and Finn, Chelsea , title =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.358454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.478163Z digest=sha256:713647228c8709380ac5f72bb7bc2db0bf98a31e5770601e28df219191b2b5da

Observation a67bb99a-ebe4-4c44-8289-3586cbede48f · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:50.483090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:06:50.483090Z digest=sha256:bb05474df8556ad795343c013079c03125bad5c1d9d5ac65acbf6d3221bcc601

Observation b159dcac-838a-448f-9c8b-ea97161450e4 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 36 (NeurIPS) , year =

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.325830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.493387Z digest=sha256:434a547d39d9ecacab8831c41d4d4c01b964cc0a12908c4a72461629a3d599c2

Observation 9a7f8a6a-1557-4061-8488-4eaf0592c0d5 · outbound

This paper cites and Smith, Laura and Kostrikov, Ilya and Levine, Sergey , title =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Smith, Laura and Kostrikov, Ilya and Levine, Sergey , title =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.276547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.499712Z digest=sha256:2d49ba11c0c1ad5da679e7876c35b090570a6585cccd6b595bb9500cc3afba13

Observation 174d633a-e74d-4d99-9707-6fb712d0387c · outbound

This paper cites and Bellemare, Marc G.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning and Bellemare, Marc G

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.247226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.505717Z digest=sha256:0cf5a86d334947281816302180f803b18fb7a136112262a3a228e2c86bcc0e75

Observation 6b04cae4-6669-4745-a134-b74f16b7b5cc · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.207655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.511800Z digest=sha256:98a54bf320f8f4d497dc3ec41c85ef4c0e3e697b679b8525a04a0b0e026b5eef

Observation 4e9b1ee6-c62b-4dea-acdb-6b592f06a8f2 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , volume =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.156893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.517833Z digest=sha256:a3f38a986d06d1d325a48661b3d22609dc63c84e600819eb2b40d493c6914295

Observation ab82142d-8529-443d-84e3-95f35dbfd25b · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning International Conference on Learning Representations (ICLR) , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.120731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.523422Z digest=sha256:9c7325a88be06aba3129a778ec4a630ed35642666020f8b52310e136060b2819

Observation 8e1bd23b-ad48-41e2-a5c0-b5d1b711de30 · outbound

This paper cites Advances in Neural Information Processing Systems 37 (NeurIPS) , year =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Advances in Neural Information Processing Systems 37 (NeurIPS) , year =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.090106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.530052Z digest=sha256:d9625c2eacab9b40634d26f1d91f115dd3cd765695644c46105ae15f356b960e

Observation f33cebe7-29a0-40cf-ac69-445cb37dc799 · outbound

This paper cites Atti del Congresso Internazionale dei Matematici, Bologna , volume =.

Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning Atti del Congresso Internazionale dei Matematici, Bologna , volume =

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:06:51.060220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T15:06:50.537588Z digest=sha256:1cdc9afc810161290e1e3f2ec0c91675e174bca7cb86238b3d60a18b2cb7dfef

Pith citing papers

No inbound Pith citation observations are available.