Pith. sign in

Paper Citation Record · LEDGER

How to Provably Improve Return Conditioned Supervised Learning?

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.08463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08463 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:10.275156Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved8
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcfc7757-4214-4f91-b109-8f106280f373 · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022.

How to Provably Improve Return Conditioned Supervised Learning? When does return-conditioned supervised learning work for offline reinforcement learning? Advances in Neural Information Processing Systems, 35:1542–1553, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:19.174686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:04.103676Z digest=sha256:fa47aa4cd0fb32fb4b7a3875d1bb2d7ec2b16a93de6807ff59b71d0b05133c52

Observation 5b9ca6b3-c313-4672-b65b-880e92d71eaa · outbound

This paper cites Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t.

How to Provably Improve Return Conditioned Supervised Learning? Decision transformer: Reinforcement learning 1For two distributionsPandQ,P≪QmeansPis absolutely continuous w.r.t

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.989128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:04.192132Z digest=sha256:713c69a9442553f1105bd21a76ea71b49e8eceef6626fb94444c2811b1ce2e86

Observation 632f3e8c-a2d3-4cd7-99cb-80ca75d930d9 · outbound

This paper cites Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,.

How to Provably Improve Return Conditioned Supervised Learning? Rvs: What is essential for offline rl via supervised learning? InInternational Conference on Learning Representations,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.778572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:04.330667Z digest=sha256:883c05262e8c5af009cb80793fe2d5778063f7e9dfc7274085ff0518738510c1

Observation 8059f672-d218-4f2a-bb9e-0686a58617fe · outbound

This paper cites Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Imitating past successes can be very suboptimal.Advances in Neural Information Processing Systems, 35:6047–6059, 2022

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.350689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:04.670336Z digest=sha256:2570101db5326dfe2928cb11c308ff32d54dc5bb755c37fc77766742d074e0cd

Observation 2ecf1c2e-1149-4524-a53e-55a7f1a8a4b7 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:04.868201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:04.868201Z digest=sha256:42fccd4bdab6b627ac01c3909e2a7c58f6722b6df3bbd5fedfb2452038011ee6

Observation 6cbb6611-c082-4463-9026-07bfd459cf24 · outbound

This paper cites A minimalist approach to offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? A minimalist approach to offline reinforcement learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:18.154249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.017626Z digest=sha256:9f14e1d3f9736d046fc59d3812f20b7b74544037a7c9bb5de0d41787bf533501

Observation 5655c842-202c-4de9-a6f0-f83ed89e05dd · outbound

This paper cites Generalized decision transformer for of- fline hindsight information matching.

How to Provably Improve Return Conditioned Supervised Learning? Generalized decision transformer for of- fline hindsight information matching

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.981506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.154444Z digest=sha256:eef83e1b45b66a229b4ad4cca499ceb8f64f673ef140706d0ee4930b97f4f3d9

Observation e2aa48c5-3aa6-4a9d-95c2-bd9a5f79d2c2 · outbound

This paper cites Act: empowering decision transformer with dynamic programming via advantage conditioning.

How to Provably Improve Return Conditioned Supervised Learning? Act: empowering decision transformer with dynamic programming via advantage conditioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.798842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.328078Z digest=sha256:ce5bbacfc56c258a0b4e9f848a71346959139e5cb3160c781726680377198448

Observation 3269d05b-354c-48d7-afa8-da6aaa6f9c35 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:05.443153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:05.443153Z digest=sha256:f21d31c6767a9fcb14de6ada1daa8905cded2f804f92c8655acd39455230a6b4

Observation c368ab03-1013-4091-9e99-89f1c5170047 · outbound

This paper cites Q-value regularized transformer for offline reinforcement learning.

How to Provably Improve Return Conditioned Supervised Learning? Q-value regularized transformer for offline reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.590710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.597675Z digest=sha256:ad9eb9123c7a2d6ec4b2de1643faa610fa586666a2d25a7baa085787cd480b8e

Observation 01b99dd9-f6c6-4d7e-8696-ab0812b89bc3 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning as one big sequence modeling problem.Advances in neural information processing systems, 34:1273– 1286, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.389293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.742356Z digest=sha256:ed198f25d7ce0521adbcee898f77076af7c4cbf52013010211a96c4a36552bd9

Observation a250da26-9d7a-4fe9-8485-5e895f9eb332 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096.

How to Provably Improve Return Conditioned Supervised Learning? Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:17.179796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:05.905336Z digest=sha256:8b986bf81857525ec51702126637654bba8982b85089656c13698570d94e97f1

Observation 7f6df767-4bcd-4739-8a4c-fab9b5d99d93 · outbound

This paper cites Reinforcement learning in robotics: A survey.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotics: A survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.979615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:06.058393Z digest=sha256:1508c809c158ce1f1e0bb18724ff810c88342c5a0ea69ee8094e115c04246784

Observation cd59261a-7881-4ac4-bc9b-a754d3a0f442 · outbound

This paper cites Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001.

How to Provably Improve Return Conditioned Supervised Learning? Quantile regression.Journal of economic perspectives, 15(4):143–156, 2001

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.814473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:06.236473Z digest=sha256:7959e9b74686ba475debc75814684b9b23e0a4ad79187eaf3f10bd11032892c8

Observation f822402f-8bc0-4912-83f8-edbe38a56fe4 · outbound

This paper cites Offline reinforcement learning with implicit q-learning.

How to Provably Improve Return Conditioned Supervised Learning? Offline reinforcement learning with implicit q-learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.614230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:06.387415Z digest=sha256:fc10e758d563633bb2f3b49262d75d91222f37f2b9fb00e9ca7a2232d0ed0bfe

Observation 6c5875f2-64b1-4600-ae59-09958a242422 · outbound

This paper cites Reward-Conditioned Policies.

How to Provably Improve Return Conditioned Supervised Learning? Reward-Conditioned Policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.527500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.527500Z digest=sha256:250869a73a87e49d9a3c6aca9446f3c57e8f8694330f65dfe5a2e68fccbb9f27

Observation b8b33b93-0360-41c1-ae74-3d9cf32ad1c5 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020.

How to Provably Improve Return Conditioned Supervised Learning? Conservative q-learning for offline reinforcement learning.Advances in Neural Information Processing Systems, 33: 1179–1191, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.417121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:06.681264Z digest=sha256:7f509c9c4ed19d29715a6cb4d3bff28d707cc6b9375bebab0818c7a320a7af5d

Observation 767bf042-42de-4d6b-abd8-08006375deda · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

How to Provably Improve Return Conditioned Supervised Learning? Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.860433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.860433Z digest=sha256:277b6b051f96e6a30937d77db6a39ae4b8f2592a2dee272a0a2224e985e708ed

Observation 2fe8ce3c-613d-4c92-9751-86e2c1fe269f · outbound

This paper cites Deep rein- forcement learning for dynamic treatment regimes on medical registry data.

How to Provably Improve Return Conditioned Supervised Learning? Deep rein- forcement learning for dynamic treatment regimes on medical registry data

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.248503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.000640Z digest=sha256:5f397f9360df48677b5ff2afc35b7ebb11a742db262febec9c5f37f2afacd5d3

Observation 8858455d-b390-4bb0-acae-6186137f7915 · outbound

This paper cites Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023.

How to Provably Improve Return Conditioned Supervised Learning? Deep spatial q- learning for infectious disease control.Journal of Agricultural, Biological and Environmental Statistics, 28(4):749–773, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:16.058975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.146714Z digest=sha256:6b4d8c1d9886156c6e66cfc30cfc1f59d8a34b4ba10adf6218c3bf87dfed9f32

Observation 9b0d61b0-f99f-49c0-a347-8858d116fba9 · outbound

This paper cites Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015.

How to Provably Improve Return Conditioned Supervised Learning? Human-level control through deep reinforcement learning.nature, 518(7540):529–533, 2015

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:15.881289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.322593Z digest=sha256:271e9d31686a443ca4176e786293dc4ce7af987117fc2498b23e6a2c12c992b1

Observation be7d3c4b-a6ba-4257-915f-69937ccbad93 · outbound

This paper cites Asymmetric least squares estimation and testing.

How to Provably Improve Return Conditioned Supervised Learning? Asymmetric least squares estimation and testing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.905451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.467721Z digest=sha256:13f6187bdc5426bf79a7c2be11af20f5329b4bedb217b564fc86adabea30f44f

Observation 90b11ce6-50d2-4a37-b5d2-e924eca9d161 · outbound

This paper cites You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022.

How to Provably Improve Return Conditioned Supervised Learning? You can’t count on luck: Why decision transformers and rvs fail in stochastic environments.Advances in neural information processing systems, 35:38966–38979, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.732547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.618289Z digest=sha256:9da7428fb4d90151dc7364b015c9106ed6c84a7fbaa0f65b0503bddd6b6ddb06

Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:07.756599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:07.756599Z digest=sha256:0365370e9093a58ea3d484cdb2858eba527893e8a049564051cbacffaf777d75

Observation 1cca5d79-1339-4d7a-a4c9-664a40f50954 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 55(2):945–990, 2022

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.536857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:07.874637Z digest=sha256:ad2b3729157ebf84fdd54f6de76e598ae30bb56e6b53abe5c43e93f2974413f5

Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.014844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.014844Z digest=sha256:5688dbfe763b32a19081cdfab913d29634d5995d47226210e2ce9cd41970e450

Observation 2afd2b9d-7214-4a41-9df1-ea4084ea1283 · outbound

This paper cites MIT press,.

How to Provably Improve Return Conditioned Supervised Learning? MIT press,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.380306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:08.194854Z digest=sha256:f1ded9e1a991e932920a9b308a481be8a863ad20b8e7a57dff33710140c6c5da

Observation 721c7b5e-3a48-45fe-9077-1e0d3ef24a7a · outbound

This paper cites Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995.

How to Provably Improve Return Conditioned Supervised Learning? Temporal difference learning and td-gammon.Communications of the ACM, 38(3):58–68, 1995

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.180697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:08.280991Z digest=sha256:e920b52b99dbcac30606b5ac00804004e5815646c9a4ff1a68a4b1dde208b4a3

Observation 87aa220a-a2ed-477e-a513-63bba5759d4e · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

How to Provably Improve Return Conditioned Supervised Learning? Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:13.005163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:08.452464Z digest=sha256:00fb84917aafec48a16d662d6fd1c2d6840f800edb4ab7fb88ddf71e71eb65d7

Observation 6141da23-8251-45f2-a0d1-7b5dbb5fea7e · outbound

This paper cites Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Return augmented decision transformer for off-dynamics reinforcement learning.arXiv preprint arXiv:2410.23450, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.615517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.615517Z digest=sha256:8b9aadfabf8973504db9feb9a7c113b0b43b021a09d16ee7213e95e19a675e3f

Observation ff035164-e227-4e2b-a65f-8cbc31daef0d · outbound

This paper cites Q-learning.Machine learning, 8:279–292, 1992.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning.Machine learning, 8:279–292, 1992

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.809949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:08.701647Z digest=sha256:1dc15d580b7d2102cbbfb5910f600dbdeec72dfcb1d85e32dd3caaf3db7f4120

Observation 53389b36-65f8-4b59-9e7f-1c542c47ee22 · outbound

This paper cites Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024.

How to Provably Improve Return Conditioned Supervised Learning? Elastic decision transformer.Advances in Neural Information Processing Systems, 36, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.608655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:08.859217Z digest=sha256:ae7dbe37d2499f7d0ec4428d81d110f661229fbd707ef2261df3734097ee95bf

Observation c631914c-ae8a-4a8d-9b28-261d8d59acbb · outbound

This paper cites A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022.

How to Provably Improve Return Conditioned Supervised Learning? A policy-guided imitation approach for offline reinforcement learning.Advances in neural information processing systems, 35: 4085–4098, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.409417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.061745Z digest=sha256:f7ae46794177d0c2e86f502b4bef68c6be898e568f04f8ccb66b00e62fb018a1

Observation f7854296-a476-4262-841e-cfb35771bd49 · outbound

This paper cites Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.219928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.142012Z digest=sha256:bc778cbd18d75c4dc126b06c99f9bc1461944fcb963544f350bcf50b97430f54

Observation 43b19276-ba88-4e84-8dcb-7e19465c5716 · outbound

This paper cites Dichotomy of control: Sepa- rating what you can control from what you cannot.

How to Provably Improve Return Conditioned Supervised Learning? Dichotomy of control: Sepa- rating what you can control from what you cannot

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:12.062449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.297843Z digest=sha256:b8810a8bb72677f6f0b60971c526334e92a8660d89177a336370d3fe6ab32c5c

Observation a50769dc-a148-47b9-9a8a-774bd96bf1c9 · outbound

This paper cites Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021.

How to Provably Improve Return Conditioned Supervised Learning? Reinforcement learning in healthcare: A survey.ACM Computing Surveys (CSUR), 55(1):1–36, 2021

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.895118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.413750Z digest=sha256:0427868f1b97e3fa0ce8a0dd953e98fa116e15c49d6dd71ea72ae62e1cf6650b

Observation 0b35ec78-9d8d-40d5-bf81-1186ecea1ed2 · outbound

This paper cites Online decision transformer.

How to Provably Improve Return Conditioned Supervised Learning? Online decision transformer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.717619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.528698Z digest=sha256:7221faa4287d8c4b3f1404a2b22d826e33d28b9c5ccb46519dd17aa4c96bc0a6

Observation 544670a7-2e87-4095-8ce6-248819a8584e · outbound

This paper cites How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266.

How to Provably Improve Return Conditioned Supervised Learning? How does goal relabeling improve sample efficiency? InProceedings of the 41st International Conference on Machine Learning, pages 61246–61266

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.514778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.716896Z digest=sha256:79d835ece427780b5a6895f383f58c31166e49b9501353619d7371e81813b9d0

Observation 77d681dd-db1b-4cca-95f7-8e794631c989 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:19:11.309789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.811315Z digest=sha256:ace5ac46737fad432dc8a27f41f32276bbcbd76981cc617dddee4fe625a97b8d

Observation 496aea22-1320-433e-a1cb-28b8a0d6fb68 · outbound

This paper cites Reinformer: Max- return sequence modeling for offline rl.

How to Provably Improve Return Conditioned Supervised Learning? Reinformer: Max- return sequence modeling for offline rl

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:11.115470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:09.935960Z digest=sha256:ba6b9b63cae5b84d55db4f01ec39c22653690981cc96f930ecb1adf4218a55c1

Observation 55a97671-7fdc-4e24-9518-930d3c5a772d · outbound

This paper cites Starting from the second stage, π⋆ starts stitching the performance of different trajectories.

How to Provably Improve Return Conditioned Supervised Learning? Starting from the second stage, π⋆ starts stitching the performance of different trajectories

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:10.095418Z digest=sha256:f1349d8bbd8f01785f1f3be117ae9462e69abd73274d60c5b4c420bbbf09f983

Observation cd6d66ec-ab1a-41b8-a2fb-2b6f7c8dac22 · outbound

This paper cites We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))].

How to Provably Improve Return Conditioned Supervised Learning? We formulate the loss function as LDT−R 2CSL =E τ [−logπ θ(·|τ)−λH(π θ(·|τ))]

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:19:10.716042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:10.275156Z digest=sha256:10b2f14749b88d949b0694b98df8e51139251fd073342d6b16103fe960d0ceee

Observation da474c82-2cbc-41c7-847d-50e88d987a07 · outbound

This paper cites an unresolved cited work.

How to Provably Improve Return Conditioned Supervised Learning? Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:19:18.566541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:19:04.464264Z digest=sha256:df9674a893327ed18ccbc219b940eefdcb2637403fc69ae6111fffffd91453d2

Pith citing papers

No inbound Pith citation observations are available.