Pith. sign in

Paper Citation Record · LEDGER

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2506.23626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23626 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:13.953628Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f13b1593-9979-4bf9-af18-bb1a1b7a84c6 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.571140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.179049Z digest=sha256:046c0d358711073927ab825f8e645d6280804d47885f533283b6278fab723131

Observation b3449306-4e61-4056-9538-1e58989b0715 · outbound

This paper cites Instructions.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:16.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.299088Z digest=sha256:13c9a018f3d0016d6f52f236c6f2c2fab903c4c6252993f340d74a4d1b56a3e7

Observation 3a6e5395-af65-4535-bcb8-925621b60483 · outbound

This paper cites • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • Adjust values to encourage the agent to complete the goal properly rather than remaining close to it

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.765168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.519280Z digest=sha256:2ba49e1e14e24b058fbfd91bd8b30b87390178300d39c409304ac09a2f7fef54

Observation 02429b1e-248e-4587-8aeb-b749060398cd · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:15.533487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.606616Z digest=sha256:9f2688a6246790db20415e8bef6a816d125f12faf7ca572589d3033e3c6eaf91

Observation 8cd012bf-5014-47db-a29b-7e198e3d1ee5 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:16.179890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.394406Z digest=sha256:286edd57137374c9c24b1af1bd3e1326e7949dad76b65bec595405393e13abbf

Observation ea1367b6-e3a6-4383-8564-c12a9ee351f8 · outbound

This paper cites • These parameters control the agent’s learning and behavior.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • These parameters control the agent’s learning and behavior

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.991458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.462055Z digest=sha256:cd253ee68c0a05e01fe603d2eba65940adb2ec14dfb713444b76d6e0eb73f83f

Observation 3f643437-5a2b-43bd-8a36-28e141c0720b · outbound

This paper cites For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games For instance, if theProblem description is talking about making the agents collide less, but does not refer hitting the fence,do not change the fence collision penalty

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:15.336253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.718492Z digest=sha256:09090c7c953fc9cbbe3706ee312d6abc7d150be55453fc23c2635268ec14ebfa

Observation d9417170-972f-4942-b930-3069e9e49353 · outbound

This paper cites • The format of the outputmust be identicalto the .txt file.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games • The format of the outputmust be identicalto the .txt file

Reference 10

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:41:15.161476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.993083Z digest=sha256:470b8e946082aa6566c754d81f8cccbcfba9f3392d31d547ed06adf3fed36815

Observation 950e8afd-d779-430d-98f7-bdd64b67600c · outbound

This paper cites in the least amount of time steps.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games in the least amount of time steps

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.974683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:12.040776Z digest=sha256:b98bf10d1c91a6c32128b81d7175249e58544cd0d2fca60ac7d816b2034f2232

Observation 9e02ff12-8e3c-4c01-a2f2-edbacc489886 · outbound

This paper cites an unresolved cited work.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:14.763880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:13.661770Z digest=sha256:0dc086e513fb92c67e8c64ee7cbefd28e22e41e754173d0ab33cfe8da6138057

Observation 59b48f51-89e1-4354-b052-6931d15109b6 · outbound

This paper cites Try to be creative with the solution, ie.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Try to be creative with the solution, ie

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.577985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:13.861797Z digest=sha256:c3f6202de6d35d44356e3c6586cb5f273c8e8086d79c9ae28e00553487dfbee0

Observation 2f60979a-0dca-4fcd-9b60-c098dc3bde44 · outbound

This paper cites Your output should be the updated reward function file only.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Your output should be the updated reward function file only

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:14.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:13.953628Z digest=sha256:92895a949a6e4a273e7b1b554610a274d295dd1acba72e50c7b5778c16afaa31

Observation c3808800-ed00-446a-8938-fb0675132a38 · outbound

This paper cites Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, Jianye Hao, Zhaopeng Meng, Yang Liu, Ruimin Shen, Yingfeng Chen, and Changjie Fan

Reference 2023

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T21:41:14.245278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:10.113675Z digest=sha256:21d638ee5d45cfacf39d584add9eaeb747dafa602db594282f19961e26eca85d

Observation 1bd7b80f-7613-4173-ba96-9553b70eddc3 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

Self-correcting Reward Shaping via Language Models for Reinforcement Learning Agents in Games OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:10.068318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:10.068318Z digest=sha256:df36e8b29a0bed663d1ab31d0c9e033aa812caf47bf78974a948ff756515c69f

Pith citing papers

No inbound Pith citation observations are available.