Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:13.953628Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2506.23626.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:13.953628Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f13b1593-9979-4bf9-af18-bb1a1b7a84c6 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3449306-4e61-4056-9538-1e58989b0715 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Instructions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a6e5395-af65-4535-bcb8-925621b60483 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02429b1e-248e-4587-8aeb-b749060398cd · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cd012bf-5014-47db-a29b-7e198e3d1ee5 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea1367b6-e3a6-4383-8564-c12a9ee351f8 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • These parameters control the agent’s learning and behavior
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f643437-5a2b-43bd-8a36-28e141c0720b · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9417170-972f-4942-b930-3069e9e49353 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • The format of the outputmust be identicalto the .txt file
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 950e8afd-d779-430d-98f7-bdd64b67600c · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games in the least amount of time steps
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e02ff12-8e3c-4c01-a2f2-edbacc489886 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b48f51-89e1-4354-b052-6931d15109b6 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Try to be creative with the solution, ie
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f60979a-0dca-4fcd-9b60-c098dc3bde44 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Your output should be the updated reward function file only
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3808800-ed00-446a-8938-fb0675132a38 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bd7b80f-7613-4173-ba96-9553b70eddc3 · outbound
Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.