Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T00:19:48.053466Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2604.20627.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T00:19:48.053466Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e2ccc47-1dc8-44f4-bf9f-7eda98aeaa9c · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Option-aware temporally abstracted value for offline goal-conditioned reinforcement learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 630a03f4-edb0-4c48-ae11-d9505cb420a5 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Full Shot Predictions for the DIII-D Tokamak via Deep Recurrent Networks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 721e048f-2b44-4b89-bbe9-2ad1e1b410c5 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Temporal Difference Flows
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c66bfe18-9a1c-4107-a25b-10e7dd7c5f43 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 838fc5f5-88e3-4bcb-b5fd-9b32323775de · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e04f272-435f-4266-9128-b8bbae6c1b1d · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Gaussian Error Linear Units (GELUs)
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c2e4c50-ad5f-47de-83c6-2097faf5b1e8 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Generative Temporal Difference Learning for Infinite-Horizon Prediction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b88afb4c-37a2-4c18-a19b-ddda345e0476 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9697fd0e-9aa8-4ef5-8e08-8d09afd4f69f · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71b73c5c-1ee4-444b-8cca-8a06710a8f94 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Automatic Reward Shaping from Confounded Offline Data
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93c2582f-c78e-4842-832f-deb3e4f6e6ff · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flow Matching for Generative Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 721075b8-417a-4601-8da8-6fc2a3cbe5dd · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flow Matching Guide and Code
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56a29c79-bb3c-482c-925d-d536bda25bad · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e597636-2626-49ea-8caf-c9ffcbacd396 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3215fa87-042d-4aa9-8040-ebc2e2211928 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Learning to Shape Rewards using a Game of Two Partners
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52b4f72d-a89e-4d2c-b643-2cc6aaad9b70 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning OGBench: Benchmarking Offline Goal-Conditioned RL
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e2a2e4a-26d9-45ed-9d73-15404fe13326 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Learning goal-conditioned policies from sub-optimal offline data via metric learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 766122dd-1a85-40bc-a3d4-6c5163bbb1b5 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Semi-parametric Topological Memory for Navigation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c069c93-9628-42a0-bf15-9fb229571ed1 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Bellman diffusion models.arXiv preprint arXiv:2407.12163
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70989f6e-e878-4c1b-a5ec-6c1d25947041 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 872eb0cf-fa4a-44df-b35d-076ed3bc29ea · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc2a4c6e-4f7f-453c-89cf-8aed4763308a · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Contrastive difference predictive coding.arXiv preprint arXiv:2310.20141, 2023a
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 259cb14f-5fd5-4920-8c06-0f0295387358 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Can a MISL fly? analysis and ingredients for mutual information skill learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 388c40ac-3e7e-437c-a9ee-6c140f4adc41 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flattening hierarchies with policy bootstrapping
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 270f7ae4-d91b-451c-8b1e-fc875d5e52d9 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef940040-5477-4b6f-abca-0042b9f320cc · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2afbd84-a533-45aa-9292-360d046cdb1f · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cddef896-1a37-4cb8-9382-b8f9bd875b09 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We start from the definition of squared Wasserstein-2 distance in Sec
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc841a1e-55db-4637-af32-a94ff0bfc09c · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10c45c45-08f8-48a9-a921-269add10669e · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning So we can lower bound the entire difference: V π∗ W (s1, g)−V π∗ W (s2, g)≥γ k−1 ·0 + k−2X t=0 γt(1−γ)∆ Φ = (1−γ)∆ Φ k−2X t=0 γt = (1−γ)∆ Φ 1−γ k−1 1−γ = (1−γ k−1)∆Φ
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b39f773b-00a8-4d0c-bad3-d5d2d382020d · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Both antmaze-large-navigateandantmaze-giant-navigateare collected with noisy expert SAC policies that repeatedly move towards randomly sampled goals (Park et al., 2024a)
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8873b6f1-c9e9-45a9-9ea9-661a625fb312 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We use a dataset of raw sensor and actuator data collected from the DIII-D tokamak located in San Diego, CA, USA
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7cd042f-2e5f-4ce4-8c31-ff552912e6b6 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Each RL algorithm was evaluated based on how closely it tracks a given goal state of the plasma
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4585f1b1-d50d-422e-9d83-e4cb18f56f39 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3704338f-0131-4f86-8d8f-f08aed0d6565 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We increased the R-Net size to an MLP of 64 units (we did not see an improvement in classification accuracy for large sizes) and used a local distance thresholdτ= 10 on all tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cebef7f-36d7-44eb-8f69-eefa3872dd4f · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We provide the specific hyperparameters used for each task in Table
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 891ed8b7-99b1-4166-96c8-84111c2ecdb0 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning 3.2 and Sec
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ac5fb54-4063-4c28-8a80-89f0be19e082 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning 4.1 for next next 1M epochs, however we did not see any consistent improvements from doing this
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1b78d31-c015-4789-9914-acfa1db7bc03 · outbound
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning For default GCIQL, policy training takes 7.2 ms per iteration
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.