Pith. sign in

Paper Citation Record · LEDGER

Adaptive Reward Design for Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2412.10917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10917 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:37:31.607767Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:30:33.390291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:30:33.977487Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8b4a402-f74d-4c99-b096-e5f8f7c13658 · outbound

This paper cites Control synthesis from linear temporal logic specifications using model-free reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Control synthesis from linear temporal logic specifications using model-free reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.056244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.416772Z digest=sha256:7bb8c56b097463909495442e953862c4a320826acdd63498163989b61b14fc03

Observation 5005d4f1-7527-460b-a175-e0e51bcb431c · outbound

This paper cites OpenAI Gym.

Adaptive Reward Design for Reinforcement Learning OpenAI Gym

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.422543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.422543Z digest=sha256:12d9064999868017724f7f6981fcd4c0faaa0a0aafcf4fa19a09a123f5cfd433

Observation 50838048-cec8-4079-b4fe-6631b96abc7d · outbound

This paper cites Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.042967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.427324Z digest=sha256:951564c2143cfe5d0be669ea24466bd311429427c5777da5996bde4050d19ea5

Observation 6028a52c-44e4-4ea1-9500-054d80fe0e5a · outbound

This paper cites Learning minimally-violating continuous control for infeasible linear temporal logic specifications.

Adaptive Reward Design for Reinforcement Learning Learning minimally-violating continuous control for infeasible linear temporal logic specifications

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.029062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.433079Z digest=sha256:bc1085be7257cf4f7b3ae7137577e4e22dff85975ee4b93c04860ea0c0c2139f

Observation a6937f62-cfca-4933-a5f8-f260ab2aaa7d · outbound

This paper cites Ltl and beyond: Formal languages for reward function specification in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Ltl and beyond: Formal languages for reward function specification in reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.016510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.440473Z digest=sha256:1f9f77d0c020400548750430cf5f70e4acccad22d3b2a0293a406143ee87bd96

Observation ccd9e90b-0801-42ff-9ef7-04b9a0d11362 · outbound

This paper cites Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications.

Adaptive Reward Design for Reinforcement Learning Foundations for restraining bolts: Reinforcement learning with ltlf/ldlf restraining specifications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:32.003421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.446769Z digest=sha256:88b6622f4fd4a100e5a8874e2a69d240a28b8ead947afc7f5e95d0ca99dc2b12

Observation 3560acac-53cd-4c1e-93b4-015dc9c65932 · outbound

This paper cites From language to goals: Inverse reinforcement learning for vision-based instruction following.

Adaptive Reward Design for Reinforcement Learning From language to goals: Inverse reinforcement learning for vision-based instruction following

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.987526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.452848Z digest=sha256:b98048f632cbefe7686c7965effb531ad68d0e9410c9a23741c65ccb2a06f76c

Observation 4c33d730-ac7d-4357-8f29-f7408f49d443 · outbound

This paper cites Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning for temporal logic control synthesis with probabilistic satisfaction guarantees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.971633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.458473Z digest=sha256:c249a13303bf1f2ff7276a3158679d5c3e03efcfcfe74c282d2b68033ce60f70

Observation 28d912a4-17f9-4163-96ae-cfc3421f9ada · outbound

This paper cites Deep reinforcement learning with temporal logics.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with temporal logics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.953462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.464583Z digest=sha256:a5dab9d073fad97516c516f76fa593197e579ba2c82b12cb3fe70c26c471280e

Observation 6c8af446-f2ec-41a3-947d-e3028e860199 · outbound

This paper cites Reward machines: Exploiting reward function structure in reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Reward machines: Exploiting reward function structure in reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.931948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.472017Z digest=sha256:17f3432993190840ea6fa8dc4089eace945126926cdd178a71ea1a162b03f1ad

Observation 6a025cf4-3809-4f04-952d-b6f39dbbf243 · outbound

This paper cites Temporal-logic-based reward shaping for continuing reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning Temporal-logic-based reward shaping for continuing reinforcement learning tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.913149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.481730Z digest=sha256:93769f9f0753d2f0f656fcec1fd3ac0bb19bf92b740fd3147c25a9e68d59ea4f

Observation 55ee23ac-cc8d-4ced-833f-b07335c6ef3b · outbound

This paper cites A composable specification language for reinforcement learning tasks.

Adaptive Reward Design for Reinforcement Learning A composable specification language for reinforcement learning tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.893885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.493989Z digest=sha256:01488d6c9ee0132ce9c209eb26cd0d2620d1b7548a4e239902b220aa6ca5f39f

Observation 2e1da156-c139-4568-972f-3140e4b28b46 · outbound

This paper cites Compositional reinforcement learning from logical specifications.

Adaptive Reward Design for Reinforcement Learning Compositional reinforcement learning from logical specifications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.879116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.509070Z digest=sha256:5991894d2802a21bd406b25173afd2969d3dbe7771d67f5c9f9c006d1135508c

Observation 990b3a30-54ab-490c-8476-c80992c7bf68 · outbound

This paper cites Model checking of safety properties.

Adaptive Reward Design for Reinforcement Learning Model checking of safety properties

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.862535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.531591Z digest=sha256:4843a046e6c04295f4c2b8d6b62d2875736ee07a4bc88ba11cd6eca8bed4e353

Observation c8d86e92-4313-4978-9209-df85a676621b · outbound

This paper cites Probabilistic planning with formal performance guarantees for mobile service robots.

Adaptive Reward Design for Reinforcement Learning Probabilistic planning with formal performance guarantees for mobile service robots

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.847530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.537253Z digest=sha256:52a075e81e4095464f12396716f1c1c424831407842dba463d556e2cca19725b

Observation e8bef7a2-3962-4188-bddf-15a9f20ee938 · outbound

This paper cites Reinforcement learning with temporal logic rewards.

Adaptive Reward Design for Reinforcement Learning Reinforcement learning with temporal logic rewards

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.833861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.547114Z digest=sha256:53dbda3e1a4e3a218771e62ddaf401a5a1967f0c4a047f61b99ff2972d7b4426

Observation 4c2b5c61-fb91-46b8-9fda-e4254a406a0f · outbound

This paper cites Continuous control with deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Continuous control with deep reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.552290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.552290Z digest=sha256:4004c7ccd69645981a969caae0a5db92b7816959abb05b99c02cae84d07f4eef

Observation d8fad458-da6e-45ce-b701-5681eb729509 · outbound

This paper cites Human-level control through deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Human-level control through deep reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.556926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.556926Z digest=sha256:c049749d8da8322093fffecd6697c94a0c2e0a5b33af871db66e597ccf79423f

Observation 1a8b3ba6-fda5-4c02-a686-379596ee1bf7 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.797080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.561606Z digest=sha256:1fcc8e9e77b46eae9dba205fee504f646b5924f8449ae6185072f5469ac7d0b9

Observation c0e50507-434f-4da3-b512-3c9c80236a2b · outbound

This paper cites Algorithms for inverse reinforcement learning.

Adaptive Reward Design for Reinforcement Learning Algorithms for inverse reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.778765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.570702Z digest=sha256:b37986deaea2ba5aba3946e3a1bcd83fb2fc664627e29d84d32bc42a7dd17cbf

Observation 08ec08ef-dc8d-42a8-80d9-42aabc55c140 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Adaptive Reward Design for Reinforcement Learning Policy invariance under reward transformations: Theory and application to reward shaping

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.757627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.584332Z digest=sha256:7473a8c8d0fbb21bf19dfbb569a0dedef01be2373211a112f81be47c56e596e7

Observation 38404a0c-da5b-459f-87bb-38bf87f9a166 · outbound

This paper cites The temporal semantics of concurrent programs.

Adaptive Reward Design for Reinforcement Learning The temporal semantics of concurrent programs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.733684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.588728Z digest=sha256:0e473fc145228d50eeddd7871d546e8658aa81a1d8386b307135c3aa047bd1be

Observation b79e5b12-0680-40e0-9c64-691dae414074 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Adaptive Reward Design for Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.717040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.593564Z digest=sha256:901c73904b5b3ed66ee761b1e57ea0d6381a614dfd350061d8103ef749bb5628

Observation 74e00a7c-6247-4e1e-9cee-cc5afc606738 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Adaptive Reward Design for Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.597864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.597864Z digest=sha256:002ccc149176cd85c46fb0822a8a36de328f8e53facd06a7557364e91890e5ee

Observation 699c7cc4-085b-4d80-aa89-8ded6c4ed214 · outbound

This paper cites Deep reinforcement learning with double q-learning.

Adaptive Reward Design for Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:37:31.701426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:37:31.602369Z digest=sha256:d7329b630cab15f1c1f271c7d6f0a2372b17b15bd55dd45206db6db091450b7c

Observation f1af1b0f-3017-46e2-9fee-6bf33ef907e7 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Adaptive Reward Design for Reinforcement Learning A survey of preference-based reinforcement learning methods

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:37:31.607767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:37:31.607767Z digest=sha256:21c321ee103d42a0664d9bc83b39ba5e1d5946c56e7e8be4f53390ec967e5f59

Pith citing papers

Observation 7def0c9b-206f-4f04-bbc5-6b2628a79826 · inbound

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning cites this paper.

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning Adaptive Reward Design for Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:30:33.983121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T00:30:33.390291Z digest=sha256:44f296235ee6b47e9e1f7d6b3b410cabaeb91c6e120e77f0d9e47197c087447a