Pith. sign in

Paper Citation Record · LEDGER

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.08158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08158 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:30:33.438380Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 784c819c-c7f1-450e-ab28-661003ee0b00 · outbound

This paper cites Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:30:34.194365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.348443Z digest=sha256:69377f78fc2eaccb413729724651f1ca2213011a1d682a1eb3eb4702bc365b56

Observation 621c1931-3a5a-4522-b387-072f2ac0beea · outbound

This paper cites Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang, Yingfeng Chen, Jianye Hao, Feng Wu, and Changjie Fan.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Yujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang, Yingfeng Chen, Jianye Hao, Feng Wu, and Changjie Fan

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.380419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.380419Z digest=sha256:5ccfcf1496f8a91d0f1e2bcb724f88e57554b502563b0b394267ed851b62b237

Observation 7def0c9b-206f-4f04-bbc5-6b2628a79826 · outbound

This paper cites Adaptive Reward Design for Reinforcement Learning.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Adaptive Reward Design for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.983121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.390291Z digest=sha256:341fe4ffec0afcd27ee911545743d528bdd39c949aca171f9597b1e043b3dae3

Observation b7ba16a5-905d-445e-8ee6-73f009a8f600 · outbound

This paper cites Improving the Effectiveness of Potential-Based Reward Shaping in Reinforcement Learning.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Improving the Effectiveness of Potential-Based Reward Shaping in Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.920720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.405381Z digest=sha256:e64aade2cbb3e281d1757faa7ff55b5394e2b384e8fa63f0e91d955dba50cb7d

Observation 0959321e-ffc7-4aa6-a979-88646145f898 · outbound

This paper cites Automating Potential-based Reward Shaping with Vision Language Model Guidance.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Automating Potential-based Reward Shaping with Vision Language Model Guidance

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.897433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.412241Z digest=sha256:2646d8f768c7090a397d83240ac791c5af32eafcdd40a0359c4fe9c978c16dfb

Observation 579c68f3-b3bb-4799-9081-306ad7477257 · outbound

This paper cites Offline Reinforcement Learning with Imputed Rewards.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Offline Reinforcement Learning with Imputed Rewards

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.805757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.423517Z digest=sha256:4370c069940001fc014321e7a8b9e98721835b9d152807f2a405165ae0f40750

Observation d28876a2-e138-4536-9f91-996396d4226d · outbound

This paper cites Training Language Models with Language Feedback.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Training Language Models with Language Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.428381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.428381Z digest=sha256:59588e4342373627594f63ef5b76543fe1d9ad9e438539aca35429c5680da073

Observation e543a49e-7729-41aa-b6ff-7b9baa04723d · outbound

This paper cites Richard S.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Richard S

Reference 20

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T00:30:33.765031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.433432Z digest=sha256:caedfed016bab64672aec5a4ca8ea9c5b8bececf189f938be1305c6a1ea812ba

Observation e4585cba-288c-442f-b3ed-abab54946091 · outbound

This paper cites Preprint: arXiv:2503.15724.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Preprint: arXiv:2503.15724

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.438380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.438380Z digest=sha256:2adce24ce1752f248daf10384edad9a6e26f997295b4fbf03cf2ae5ccecbe756

Observation 0ac6bbba-f328-4760-b7dc-0a49d7ae0571 · outbound

This paper cites SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning

Reference 2003

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.959833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.395422Z digest=sha256:25b2dfc01caa94d3649c624d06b0fe20b594679a6dafa86b00b071cda9aa1d41

Observation 12c4f907-4d91-4413-a2da-1075344c9f3c · outbound

This paper cites Concrete Problems in AI Safety.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Concrete Problems in AI Safety

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.337729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.337729Z digest=sha256:7851eb8926c970fddd7fd88684f9b15af56ef28d485c91fa870645d3e7374660

Observation 06a4be0d-ce7c-44bc-8c1e-b0e5d628c28c · outbound

This paper cites Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Unpacking Reward Shaping: Understanding the Benefits of Reward Engineering on Sample Complexity

Reference 2010

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:30:34.023186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.374603Z digest=sha256:94288c71f5de9c3c13fd7f6c59e604d3a408eca42713ba1d47c127fd8bd5ef20

Observation 4e5647f9-ed4c-4576-a22c-724cd380c882 · outbound

This paper cites Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Extracting Heuristics from Large Language Models for Reward Shaping in Reinforcement Learning

Reference 2016

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:34.061337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.364295Z digest=sha256:f61618b0694c5060507e9a095b3e8e34d01ab0f87c7e4bda791268d5ea9b027e

Observation fccf0884-caea-4685-8570-debabaf3e1fa · outbound

This paper cites Vision-Language Models as a Source of Rewards.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Vision-Language Models as a Source of Rewards

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.353399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.353399Z digest=sha256:362439823a6df6a88123ad6819bfd5be2df1265a9f891f543e0329a0bb7a9cc4

Observation a9156aeb-e0e3-4b49-9afb-f427b7132c33 · outbound

This paper cites Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Potential-Based Intrinsic Motivation: Preserving Optimality With Complex, Non-Markovian Shaping Rewards

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.369554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.369554Z digest=sha256:9800b3d8079ebd1a59790f8082395671b7395bd4ecb5641d9141186158b43df9

Observation 7cd305ff-1310-4359-90b9-744a94f37100 · outbound

This paper cites Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.385098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.385098Z digest=sha256:1275cd8ad9507059908a462548d05f8a76f9456f4764bc45554dde557fdc0c54

Observation 53f5622c-5e08-43f4-855e-6d73c66e15bc · outbound

This paper cites Preprint: arXiv:2104.06411.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Preprint: arXiv:2104.06411

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.418659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.418659Z digest=sha256:941a4466f5f3b70f725d07fe96dc4f55e7101db98348d2f895fdbdc3dee5d1ba

Observation 8d706fa9-bb33-4e1c-881f-420315478d9e · outbound

This paper cites 2022.1027340.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning 2022.1027340

Reference 2022

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:30:33.343177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.343177Z digest=sha256:3673a0c91054f27b1844b1c7d569084fa96816c2574bee1be63d3a00f33085b7

Observation dfafb29f-ee29-4220-a9a3-2eca840c1fcf · outbound

This paper cites InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:30:33.400576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:30:33.400576Z digest=sha256:47e1dc8b5f15048ac0c9cd6aab657f8983138d8861371c81964baae7869dd034

Observation eb2dda5a-b769-4883-9785-8232e37977b4 · outbound

This paper cites Useful Policy Invariant Shaping from Arbitrary Advice.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Useful Policy Invariant Shaping from Arbitrary Advice

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:30:34.083387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.358818Z digest=sha256:ecfee7cc61521794338c93e306be2bd46fafece1580ec3dac2a7be010298db8c

Observation 59c866d2-a085-43a9-be54-9e9d869c7020 · outbound

This paper cites AdrianK.AgoginoandKaganTumer.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning AdrianK.AgoginoandKaganTumer

Reference 2025

Resolution
verified exact
doi, observed 2026-08-12T00:30:33.487245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:30:33.332256Z digest=sha256:9412ed118849be9594e5de7630439f9dbaafe6c519b6b682450b4d7bce272f39

Pith citing papers

No inbound Pith citation observations are available.