Pith. sign in

Paper Citation Record · LEDGER

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2604.20627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20627 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:19:48.053466Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact23
  • verified fuzzy10
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e2ccc47-1dc8-44f4-bf9f-7eda98aeaa9c · outbound

This paper cites Option-aware temporally abstracted value for offline goal-conditioned reinforcement learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Option-aware temporally abstracted value for offline goal-conditioned reinforcement learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.170209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:90807c0144af81249a9f73c2b38e8f297730ce19e426d346018d5b24c64c8bc5

Observation 630a03f4-edb0-4c48-ae11-d9505cb420a5 · outbound

This paper cites Full Shot Predictions for the DIII-D Tokamak via Deep Recurrent Networks.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Full Shot Predictions for the DIII-D Tokamak via Deep Recurrent Networks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.167457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:7cf5934aad8f185a009aef9d9261c758d8f19af9a5fc8cd861d3eaa73f9004f2

Observation 721e048f-2b44-4b89-bbe9-2ad1e1b410c5 · outbound

This paper cites Temporal Difference Flows.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Temporal Difference Flows

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.158900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:9731a494b161acf0ab1d186eb2da8d030f5af2ee2b51b4b7462a9db0f3ab77eb

Observation c66bfe18-9a1c-4107-a25b-10e7dd7c5f43 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:19:17.477461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:1015054c6655933c06c4fab71d490515a139f8ea32f64937982004d202d5306d

Observation 838fc5f5-88e3-4bcb-b5fd-9b32323775de · outbound

This paper cites Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.189847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:098f02d785c185aa34cef3c27ff41b68695b931d92467019d45d23a10e4aaa41

Observation 8e04f272-435f-4266-9128-b8bbae6c1b1d · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.184872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:83afbe3208b76212a3121f46809b6f28f63873dfbb624c759a57332804b0687c

Observation 5c2e4c50-ad5f-47de-83c6-2097faf5b1e8 · outbound

This paper cites Generative Temporal Difference Learning for Infinite-Horizon Prediction.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Generative Temporal Difference Learning for Infinite-Horizon Prediction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.173297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:6bcd42cf5a91fb2eb6b13022444d10f64cd790a57266962b4b1856b066f169dc

Observation b88afb4c-37a2-4c18-a19b-ddda345e0476 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T00:24:47.182085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:b9ce26794c82f3f3548e11155d30eb3bf2aec59e018965bbb14d68f45dac05e9

Observation 9697fd0e-9aa8-4ef5-8e08-8d09afd4f69f · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:47:05.700994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:a2ab32d52399a54a81ab159406edb60204395d70c5ca86a50eefb4abb294c57d

Observation 71b73c5c-1ee4-444b-8cca-8a06710a8f94 · outbound

This paper cites Automatic Reward Shaping from Confounded Offline Data.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Automatic Reward Shaping from Confounded Offline Data

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.192353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:076bc273ded89365f8691099485574e12d9e1bd304a25aac9c7f3a8fd940854b

Observation 93c2582f-c78e-4842-832f-deb3e4f6e6ff · outbound

This paper cites Flow Matching for Generative Modeling.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flow Matching for Generative Modeling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T00:24:47.194847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:6daba442e1a684ff79aa1787509aecdb90b11f032f70492f0e4d166c7872c998

Observation 721075b8-417a-4601-8da8-6fc2a3cbe5dd · outbound

This paper cites Flow Matching Guide and Code.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flow Matching Guide and Code

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:28:14.320594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:85896a3f7b6fb8a64bd586fc58f23405b7b5500c342668002b457fef20933268

Observation 56a29c79-bb3c-482c-925d-d536bda25bad · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.179747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:2088bfc2ec57c96c502075fe6c0e78931df6a11aec27b92d91175d0eea718127

Observation 3e597636-2626-49ea-8caf-c9ffcbacd396 · outbound

This paper cites Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Highly Efficient Self-Adaptive Reward Shaping for Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.176236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:c3c2ef39f88d473f8b047ac99b865bb6fec1207d0d429ab8b049c3297824ec62

Observation 3215fa87-042d-4aa9-8040-ebc2e2211928 · outbound

This paper cites Learning to Shape Rewards using a Game of Two Partners.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Learning to Shape Rewards using a Game of Two Partners

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.226301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:fe638552ef7713a27f7b452fa15be431bd32bf67f88d99f831f0ac884dea2ddd

Observation 52b4f72d-a89e-4d2c-b643-2cc6aaad9b70 · outbound

This paper cites OGBench: Benchmarking Offline Goal-Conditioned RL.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning OGBench: Benchmarking Offline Goal-Conditioned RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.205459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:943a7693535a4d4b5c50f8f21d100e3e8b4eff0c90b4a37d1820836de1f2decc

Observation 7e2a2e4a-26d9-45ed-9d73-15404fe13326 · outbound

This paper cites Learning goal-conditioned policies from sub-optimal offline data via metric learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Learning goal-conditioned policies from sub-optimal offline data via metric learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.210600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:a63ae3e1fd48a165add20654438411cd618086cbfda82581c2cb92f47e7fabeb

Observation 766122dd-1a85-40bc-a3d4-6c5163bbb1b5 · outbound

This paper cites Semi-parametric Topological Memory for Navigation.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Semi-parametric Topological Memory for Navigation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.213063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:92643ba59ea9ea5b789ff767137c06669c3de1c225706119ed610edc90fa66e8

Observation 1c069c93-9628-42a0-bf15-9fb229571ed1 · outbound

This paper cites Bellman diffusion models.arXiv preprint arXiv:2407.12163.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Bellman diffusion models.arXiv preprint arXiv:2407.12163

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.218256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:37680ccf1ec07cc3afbc02c741e8a9eef1212eedbf3dd301110401cf3a43be0b

Observation 70989f6e-e878-4c1b-a5ec-6c1d25947041 · outbound

This paper cites Dual RL: Unification and New Methods for Reinforcement and Imitation Learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Dual RL: Unification and New Methods for Reinforcement and Imitation Learning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.200049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:5b8b05dcbff5da36f63760b7f2be5765b27442808153cf68324cf24e30b6ce75

Observation 872eb0cf-fa4a-44df-b35d-076ed3bc29ea · outbound

This paper cites Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.220742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:28cc22d2595b02e0f82ae2460947c7620272d71bf756ae6ace57e144f27b2c41

Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · outbound

This paper cites Reward Models in Deep Reinforcement Learning: A Survey.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.223690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:270e71076527435f39bc403edb44399677a0655913052183fea1acbd5a27d076

Observation fc2a4c6e-4f7f-453c-89cf-8aed4763308a · outbound

This paper cites Contrastive difference predictive coding.arXiv preprint arXiv:2310.20141, 2023a.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Contrastive difference predictive coding.arXiv preprint arXiv:2310.20141, 2023a

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.197432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:4dc1ae878833af506e4c32d3dcfa49f9145b69cf094c2cb02d6899343bd5cb4a

Observation 259cb14f-5fd5-4920-8c06-0f0295387358 · outbound

This paper cites Can a MISL fly? analysis and ingredients for mutual information skill learning.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Can a MISL fly? analysis and ingredients for mutual information skill learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.207972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:75e7c75f43231bd0ba363af766dc34fd669971e756bc91e4bcdf14f8c7fba20c

Observation 388c40ac-3e7e-437c-a9ee-6c140f4adc41 · outbound

This paper cites Flattening hierarchies with policy bootstrapping.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Flattening hierarchies with policy bootstrapping

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.202852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:83d8d1ff799649462928bc8abdc7eb15e5535d5d677675572f89986ce813dddb

Observation 270f7ae4-d91b-451c-8b1e-fc875d5e52d9 · outbound

This paper cites an unresolved cited work.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-23T11:27:55.427758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:00962d02a1b6e846c6c6dbb795bada301a7e51a57e4eec519c152440ef3f4b41

Observation ef940040-5477-4b6f-abca-0042b9f320cc · outbound

This paper cites an unresolved cited work.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-23T11:27:55.420385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:550de074c588eca637ba25dc6c7ae15840d03475b4bce8892dabd9d945d30c21

Observation d2afbd84-a533-45aa-9292-360d046cdb1f · outbound

This paper cites an unresolved cited work.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-23T11:27:55.437349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:d64352080b026a6df61bfff07d4ef6ff0657ab111fc7b176a2212283fe8ae2af

Observation cddef896-1a37-4cb8-9382-b8f9bd875b09 · outbound

This paper cites We start from the definition of squared Wasserstein-2 distance in Sec.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We start from the definition of squared Wasserstein-2 distance in Sec

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.430729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:c22a45f19d3b67a069c37affa1ba8a086d0d5dfdeeaac06f13e56a151deb58b5

Observation bc841a1e-55db-4637-af32-a94ff0bfc09c · outbound

This paper cites an unresolved cited work.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-23T11:27:55.407343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:a8cd51da38274393f10f07925397c47acd6bfe76c3ca097fd0787546e22e4090

Observation 10c45c45-08f8-48a9-a921-269add10669e · outbound

This paper cites So we can lower bound the entire difference: V π∗ W (s1, g)−V π∗ W (s2, g)≥γ k−1 ·0 + k−2X t=0 γt(1−γ)∆ Φ = (1−γ)∆ Φ k−2X t=0 γt = (1−γ)∆ Φ 1−γ k−1 1−γ = (1−γ k−1)∆Φ.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning So we can lower bound the entire difference: V π∗ W (s1, g)−V π∗ W (s2, g)≥γ k−1 ·0 + k−2X t=0 γt(1−γ)∆ Φ = (1−γ)∆ Φ k−2X t=0 γt = (1−γ)∆ Φ 1−γ k−1 1−γ = (1−γ k−1)∆Φ

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.441151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:b423c6423d0cc69be391698f99b60b5c094cba301eb2d5d63b404dcd6762bd11

Observation b39f773b-00a8-4d0c-bad3-d5d2d382020d · outbound

This paper cites Both antmaze-large-navigateandantmaze-giant-navigateare collected with noisy expert SAC policies that repeatedly move towards randomly sampled goals (Park et al., 2024a).

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Both antmaze-large-navigateandantmaze-giant-navigateare collected with noisy expert SAC policies that repeatedly move towards randomly sampled goals (Park et al., 2024a)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.434156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:e929c0155b917c934f68676613074c1deb9ce5f02fafc502bdcabf33ceee4942

Observation 8873b6f1-c9e9-45a9-9ea9-661a625fb312 · outbound

This paper cites We use a dataset of raw sensor and actuator data collected from the DIII-D tokamak located in San Diego, CA, USA.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We use a dataset of raw sensor and actuator data collected from the DIII-D tokamak located in San Diego, CA, USA

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.448423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:c8d0c3407f02958725bc20d646abbcab1fa9aab3d8973864feabe716fddb2dab

Observation f7cd042f-2e5f-4ce4-8c31-ff552912e6b6 · outbound

This paper cites Each RL algorithm was evaluated based on how closely it tracks a given goal state of the plasma.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Each RL algorithm was evaluated based on how closely it tracks a given goal state of the plasma

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.455107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:8f89f60ab5d1f32a87fa8d5632bdc1c70a7e90488988effa5477f208f67e8ed5

Observation 4585f1b1-d50d-422e-9d83-e4cb18f56f39 · outbound

This paper cites an unresolved cited work.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-23T11:27:55.451431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:4220587c42791fb3c5bded770146ca843e0b9a6fbbe3f07f0808f2d50a74acac

Observation 3704338f-0131-4f86-8d8f-f08aed0d6565 · outbound

This paper cites We increased the R-Net size to an MLP of 64 units (we did not see an improvement in classification accuracy for large sizes) and used a local distance thresholdτ= 10 on all tasks.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We increased the R-Net size to an MLP of 64 units (we did not see an improvement in classification accuracy for large sizes) and used a local distance thresholdτ= 10 on all tasks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.424707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:255aa9ac2e5df46406198179e2801f1af5e38af1ac13740530b78b7e291e009a

Observation 4cebef7f-36d7-44eb-8f69-eefa3872dd4f · outbound

This paper cites We provide the specific hyperparameters used for each task in Table.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning We provide the specific hyperparameters used for each task in Table

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.444953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:7f014b31496a0a78a49b7cb48469efc9434e389acfa6ac44f2d4b4c48e67e24d

Observation 891ed8b7-99b1-4166-96c8-84111c2ecdb0 · outbound

This paper cites 3.2 and Sec.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning 3.2 and Sec

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.410453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:43af7b9ec6f3b88fab7b63a5466d1e0bee8418fac62839694cb12a881b1c3aff

Observation 4ac5fb54-4063-4c28-8a80-89f0be19e082 · outbound

This paper cites 4.1 for next next 1M epochs, however we did not see any consistent improvements from doing this.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning 4.1 for next next 1M epochs, however we did not see any consistent improvements from doing this

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.413770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:d7add03756109d93f64456179b9ffce21007788359bd989d9b2da8e2ac9b9385

Observation e1b78d31-c015-4789-9914-acfa1db7bc03 · outbound

This paper cites For default GCIQL, policy training takes 7.2 ms per iteration.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning For default GCIQL, policy training takes 7.2 ms per iteration

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:27:55.417304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:3215b2c25d463a2201f90fe7e22db7558250f79b222b746e85807276cb2ddd50

Pith citing papers

No inbound Pith citation observations are available.