Pith. sign in

Paper Citation Record · LEDGER

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2411.17861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17861 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:52:46.731946Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:28:06.474838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:28:07.394748Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26a01fb4-8144-45d4-b498-45d2635589e9 · outbound

This paper cites Robustness measures and monitors for time window temporal logic.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Robustness measures and monitors for time window temporal logic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.562833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.484155Z digest=sha256:a079ba7b5529bd631fae6d58cb4675775175b3dee40e8bd6c7ecaa19a076dce8

Observation b831f535-cec6-4552-8dd3-5d889cac1b20 · outbound

This paper cites Q-learning for robust satisfaction of signal temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Q-learning for robust satisfaction of signal temporal logic specifications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.535758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.493814Z digest=sha256:aa9cac30037e53a77c575686eada61adea4a320ed0c87012ab7c52f2ca100112

Observation 02c8c338-4d29-4487-baa7-c5d32e88ac3b · outbound

This paper cites u diger Ehlers, Bettina K \.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards u diger Ehlers, Bettina K \

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.502104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.501307Z digest=sha256:c15715560d358c1b0bd6c53ece66ccb32affddf452d9edfa2d03d1c8df906c0c

Observation 575564f0-34d5-4a2b-a6cb-eb8881e3f4d3 · outbound

This paper cites Temporal-logic-constrained hybrid reinforcement learning to perform optimal aerial monitoring with delivery drones.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Temporal-logic-constrained hybrid reinforcement learning to perform optimal aerial monitoring with delivery drones

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.478714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.512189Z digest=sha256:95900523c58bddf620681e8d95101e8dde32bb61f12ed5b41e32f03573d287bc

Observation 4f82c5e7-bdc9-4640-ba75-73677b8c7b4c · outbound

This paper cites Principles of model checking.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Principles of model checking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.455806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.532382Z digest=sha256:52dfbef42f60050ee01425bd831e53fd4fa3bc4766fce59f59fce454b88bee3f

Observation 46279274-41d4-4a85-be1c-63164cd10115 · outbound

This paper cites Structured reward shaping using signal temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Structured reward shaping using signal temporal logic specifications

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.431795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.554330Z digest=sha256:fe0db0021f6511d3d56c79a9b2a9c769503fd8c87174fa0f3beddfca92a41344

Observation 507a73d9-7d7b-45a8-961e-1653a957f7cb · outbound

This paper cites Efficient online reinforcement learning with offline data.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Efficient online reinforcement learning with offline data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.561945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.561945Z digest=sha256:55a20a999e355a62db3f4a44e98764110b8074ddfd29ed1344e1be89cc913fd9

Observation 140370b8-4ab6-4a8c-bfd8-1f734537dc12 · outbound

This paper cites Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.392580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.571214Z digest=sha256:e6cdd924df7abc18709c21881b7febe2cc6244d10514ce7fc1b57871dd6f3469

Observation 4f49832c-82b9-4a68-8481-deeebed5c022 · outbound

This paper cites Provably efficient exploration in policy optimization.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Provably efficient exploration in policy optimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.365131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.579436Z digest=sha256:25382f3c3e9109f54d80507c66c6fbf4e6afb66becc3e76d9caca00c7c7e84f2

Observation 074d7896-4186-4531-a08c-f8cb89d001db · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Decision transformer: Reinforcement learning via sequence modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.586476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.586476Z digest=sha256:b5e51081742226f1e263416bbf33ab48d509a25b9c611b43d367ff0db9cc4b34

Observation a71427fb-82c6-4d19-a888-d51377e037bc · outbound

This paper cites Mirror learning: A unifying framework of policy optimisation.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Mirror learning: A unifying framework of policy optimisation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.325509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.591702Z digest=sha256:3b98e2717dc2fc1a0655abe2774c9684c8d6430bd905d00df833de436e63c775

Observation a537d976-ec31-47ec-a8b5-e08437152fc5 · outbound

This paper cites Imitation Bootstrapped Reinforcement Learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Imitation Bootstrapped Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.600129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.600129Z digest=sha256:71cef1788b0b82a2ca9e050048822ce0e7d728c92c9c7e454739da7f8e9b4e2e

Observation 315a9cc6-2786-4d74-b15d-c7d494085f9f · outbound

This paper cites A tutorial on mm algorithms.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards A tutorial on mm algorithms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.300989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.606014Z digest=sha256:1fbed3b9166d9e233f1933ed04da3e448f4a3a47db2dd472362411a293fc5902

Observation abb12aef-0a87-4c8d-8847-37421b20b81b · outbound

This paper cites Reward machines: Exploiting reward function structure in reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reward machines: Exploiting reward function structure in reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.612212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.612212Z digest=sha256:87b2cc11dd2f03ee4f1fe4e326e69d956c126c3b7e031af77c17408e5423cb45

Observation 19539de4-6464-4d0d-917a-110a92845c3b · outbound

This paper cites Approximately optimal approximate reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Approximately optimal approximate reinforcement learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.617397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.617397Z digest=sha256:da8e1c29ff0ca41fbb3e00a04a8077deced72b2858dfcb607f950910d896fbb1

Observation c99c3e8f-28c8-4e46-b47c-e89d8c6d131c · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.624935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.624935Z digest=sha256:58009d424f49e4a83c73150d16077033eec28d7f61897f3aeb5591b4092984ec

Observation 35a87e66-ffdd-4134-b9d5-5cb559e10415 · outbound

This paper cites Reinforcement learning with temporal logic rewards.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning with temporal logic rewards

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.229610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.630179Z digest=sha256:8adeaecefc592d83422f3bfd9599c2e7b86076ab906314b076a4e10eac853eae

Observation 70a1a7f9-d3d9-4b9e-a01d-af8814f4e334 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.637160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.637160Z digest=sha256:061e78a4ebdd90648e3a0f7090f6275df7bc78ad5477e8f11ae3531076d3ff61

Observation 0547adb5-2f2c-43f0-a5c7-d748303d54f0 · outbound

This paper cites Advice-guided reinforcement learning in a non-markovian environment.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Advice-guided reinforcement learning in a non-markovian environment

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.193723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.644641Z digest=sha256:55384af25cf2d550fab8a69a11739234488ddba07f272d53e2d1300da9a69dd8

Observation ba6f8e0c-4952-4f05-bafb-8bad039f17c5 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Policy invariance under reward transformations: Theory and application to reward shaping

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.163130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.651318Z digest=sha256:6ab02cd18b074bc446cb60ac53130bcabf7c01010ac4707f6394cfdb211b653c

Observation 9675e615-6219-4b16-8b3d-31c0b35025b3 · outbound

This paper cites Implicit human perception learning in complex and unknown environments.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Implicit human perception learning in complex and unknown environments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.129012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.658404Z digest=sha256:73fb031985d72328796ba4b5e1e39343666482c84303f5b153bed472f1f534a8

Observation d61732d2-32ce-49e4-ab93-009ff1927aae · outbound

This paper cites Agnostic system identification for model-based reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Agnostic system identification for model-based reinforcement learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.106305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.663767Z digest=sha256:723ac0b105b27e86a7cd3877a11ee04548ccf381cf69f1c6ada235c9f1928467

Observation 4932ad35-cd2e-4111-a6f4-7adc7d0841c7 · outbound

This paper cites Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.083880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.669943Z digest=sha256:283051cdbd079c05fbf16e1e7bccb6cd025472237ce3f718de4ffc3f0a43c245

Observation c8228c20-aed5-48aa-86a2-167f48d9e72f · outbound

This paper cites Kickstarting Deep Reinforcement Learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Kickstarting Deep Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.677141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.677141Z digest=sha256:f1f21d907dc3750eaf1e321977ced1daa289a4261734143c777dbd6bc76e7f09

Observation 5cfa168b-e794-4500-9c49-3ef6bf4068e7 · outbound

This paper cites Trust region policy optimization.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trust region policy optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:47.060838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.683744Z digest=sha256:f1631c3c2698db08a32e0e3eb0e93bb8d86be85f49eca39642fa7ace6fb46a87

Observation c9d8d462-38ba-43fd-9fda-13a0da0b570a · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.695180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.695180Z digest=sha256:0b1c8ff71964f4dd5087d99ea5b113f0fb87cd84decb58cd4059ac52c6a50ddc

Observation 9734904f-ca20-45c8-9a93-1ba29d8df102 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.701697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.701697Z digest=sha256:5febb28d66f07f2b6d00eb503c87718c920bb78ad274a6f7312da416baec139d

Observation af193ef4-b9b9-412f-aea1-c2643a4c53c6 · outbound

This paper cites Reinforcement learning: An introduction.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning: An introduction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T11:52:46.714546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:52:46.714546Z digest=sha256:48d83d81d7d7a870d1d3f910d6c387be8db5f52cd5c6be7b42d19e8fbde3ff7f

Observation 1c0d99f7-56a0-45ca-af40-07e6a58348ca · outbound

This paper cites an unresolved cited work.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:52:47.005125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.720242Z digest=sha256:d2c22e16f7a2db726783bc595d07aaee6a7ced5786ca0aa2f86e8aec80de8926

Observation 0b05f1c8-2b1c-425d-b52e-2168edbf0018 · outbound

This paper cites Time window temporal logic.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Time window temporal logic

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:46.971335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.726539Z digest=sha256:52e446b01e3fe7d0f3ea57275efd714c145810039d402a9d563193085a5024f6

Observation 970abbdd-c3a1-4355-b40e-bdf92c9c81c6 · outbound

This paper cites Joint inference of reward machines and policies for reinforcement learning.

Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Joint inference of reward machines and policies for reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:52:46.925300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T11:52:46.731946Z digest=sha256:cbb2c3197317d76bfc42ed31cb668da1f6af5ba83f35ad272e4964be80f9ab6e

Pith citing papers

Observation 8ff85624-1885-40b4-9c98-286e164ae58a · inbound

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins cites this paper.

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:28:07.444921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:28:06.474838Z digest=sha256:bed0ef6578c1afd08af9baf9ab3a1ad9dfa7b93722f90a4f4027ed203caa239c