Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:52:46.731946Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2411.17861.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:52:46.731946Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:28:06.474838Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T20:28:07.394748Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26a01fb4-8144-45d4-b498-45d2635589e9 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Robustness measures and monitors for time window temporal logic
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b831f535-cec6-4552-8dd3-5d889cac1b20 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Q-learning for robust satisfaction of signal temporal logic specifications
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02c8c338-4d29-4487-baa7-c5d32e88ac3b · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards u diger Ehlers, Bettina K \
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 575564f0-34d5-4a2b-a6cb-eb8881e3f4d3 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Temporal-logic-constrained hybrid reinforcement learning to perform optimal aerial monitoring with delivery drones
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4f82c5e7-bdc9-4640-ba75-73677b8c7b4c · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Principles of model checking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46279274-41d4-4a85-be1c-63164cd10115 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Structured reward shaping using signal temporal logic specifications
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 507a73d9-7d7b-45a8-961e-1653a957f7cb · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Efficient online reinforcement learning with offline data
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140370b8-4ab6-4a8c-bfd8-1f734537dc12 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Overcoming exploration: Deep reinforcement learning for continuous control in cluttered environments from temporal logic specifications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4f49832c-82b9-4a68-8481-deeebed5c022 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Provably efficient exploration in policy optimization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 074d7896-4186-4531-a08c-f8cb89d001db · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Decision transformer: Reinforcement learning via sequence modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71427fb-82c6-4d19-a888-d51377e037bc · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Mirror learning: A unifying framework of policy optimisation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a537d976-ec31-47ec-a8b5-e08437152fc5 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Imitation Bootstrapped Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315a9cc6-2786-4d74-b15d-c7d494085f9f · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards A tutorial on mm algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation abb12aef-0a87-4c8d-8847-37421b20b81b · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reward machines: Exploiting reward function structure in reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19539de4-6464-4d0d-917a-110a92845c3b · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Approximately optimal approximate reinforcement learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99c3e8f-28c8-4e46-b47c-e89d8c6d131c · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Conservative q-learning for offline reinforcement learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a87e66-ffdd-4134-b9d5-5cb559e10415 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning with temporal logic rewards
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70a1a7f9-d3d9-4b9e-a01d-af8814f4e334 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0547adb5-2f2c-43f0-a5c7-d748303d54f0 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Advice-guided reinforcement learning in a non-markovian environment
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba6f8e0c-4952-4f05-bafb-8bad039f17c5 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Policy invariance under reward transformations: Theory and application to reward shaping
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9675e615-6219-4b16-8b3d-31c0b35025b3 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Implicit human perception learning in complex and unknown environments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d61732d2-32ce-49e4-ab93-009ff1927aae · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Agnostic system identification for model-based reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4932ad35-cd2e-4111-a6f4-7adc7d0841c7 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trajectron++: Dynamically-feasible trajectory forecasting with heterogeneous data
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8228c20-aed5-48aa-86a2-167f48d9e72f · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Kickstarting Deep Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cfa168b-e794-4500-9c49-3ef6bf4068e7 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Trust region policy optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c9d8d462-38ba-43fd-9fda-13a0da0b570a · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9734904f-ca20-45c8-9a93-1ba29d8df102 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af193ef4-b9b9-412f-aea1-c2643a4c53c6 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Reinforcement learning: An introduction
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0d99f7-56a0-45ca-af40-07e6a58348ca · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b05f1c8-2b1c-425d-b52e-2168edbf0018 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Time window temporal logic
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 970abbdd-c3a1-4355-b40e-bdf92c9c81c6 · outbound
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards Joint inference of reward machines and policies for reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ff85624-1885-40b4-9c98-286e164ae58a · inbound
ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.